Which of these 100 apps can you actually integrate?

An agent-built buildability assessment of 100 third-party apps — auth model, whether a developer can self-serve credentials, whether an MCP server already exists, and what actually blocks the build. Every claim carries an evidence URL and a confidence score; a 20-app verification sample puts claim accuracy at 92.5%.

100 apps · 10 categories 80 claims verified across 3 independent methods 92.5% claim accuracy Data as of 2026-08-17

1  Headline patterns

what the 100 rows add up to

Each number below is followed by what it means for someone building an integration toolkit — the interpretation is the point, not the count.

94 / 100
Auth is nearly solved
API key74
OAuth 2.059
Basic10
JWT7
No public auth5

94 of 100 apps authenticate with an API key, OAuth 2.0, or both (39 support both). Two auth paths cover the entire catalogue. Auth is not where the engineering cost lives — so don't budget for it as if it were.

35
Zero-friction easy wins

35 apps are easy + self-serve with no blocker at all — documented REST API, instant credentials, nothing to negotiate. Start here: roughly a third of the catalogue can ship without a single conversation.

30
Blocked by paperwork, not code

Of the 65 apps with a blocker, 30 are approval gates — app review, partner programmes, sales contracts, developer registration. Only 12 are architectural. The bottleneck is a partnerships motion, not engineering time.

40
First-party MCP servers
Dev infra8/10
Data & scraping7/10
Productivity6/10
Ecommerce1/10
Comms1/10
Marketing / ads1/10

40 vendors ship and maintain their own MCP server; 1 is community-only. In dev-infra and productivity the surface is already built — adopt, don't rebuild. Ecommerce, comms and ads are wide open by contrast.

20 / 20 vs 5 / 10
Category decides difficulty
Productivity10/10
Dev infra10/10
CRM9/10
Comms9/10
Support8/10
Fintech6/10
Marketing / ads5/10

Self-serve rate by category. Productivity and dev-infra are 20/20 self-serve and 18/20 easy; marketing/ads is 5/10 self-serve with 4 hard gates (Google, Meta, LinkedIn, Pinterest). Sequence the roadmap by category — and staff ads platforms with BD, not just engineers.

21
Need outreach before code

21 apps are gated, partial, or unknown-access; 10 are hard or blocked outright. Qualify these before committing sprint capacity — 2 (fanbasis, iPayX) have no public API at all and can't be built against today.

2  The 100

filter, search, sort · ✓ = independently verified

Every row is one assessed app. The confidence bar is the model's own self-reported confidence at extraction time; anything under 0.60 was routed to human review.

100 of 100
# ▲▼ App ▲▼ Category ▲▼ Auth Access ▲▼ MCP ▲▼ Verdict ▲▼ Main blocker Evidence Conf. ▲▼
1SalesforceCRM and Salesoauth2, jwtself serveunresolvedmediumExternal Client App (formerly Connected App, renamed Spring '26) + OAuth setup and per-org instance URLs add integration complexitysource0.85
2HubSpotCRM and Salesoauth2, api_keyself servefirst-partyeasynonesource0.90
3PipedriveCRM and Salesapi_key, oauth2self serveunresolvedeasynonesource0.85
4AttioCRM and Salesoauth2, api_keyself servefirst-partyeasynonesource0.80
5TwentyCRM and Salesapi_keyself servenoneeasysmaller ecosystem; API surface still evolving with the open-source projectsource0.65
6PodioCRM and Salesoauth2, api_keyself servenoneeasyproduct is in low-investment/maintenance mode, so long-term API viability is a risksource0.80
7Zoho CRMCRM and Salesoauth2self serveunresolvedmediummulti-datacenter endpoints (zoho.com/.eu/.in) and OAuth-only auth complicate a generic connectorsource0.50
8CloseCRM and Salesapi_key, oauth2self servefirst-partyeasynonesource0.70
9CopperCRM and Salesapi_key, oauth2self servenoneeasynonesource0.85
10DealCloudCRM and Salesapi_key, oauth2gatedfirst-partyhardenterprise product — API keys exist per client site, so you need a DealCloud customer instance to build/test againstsource0.70
11ZendeskSupport and Helpdeskoauth2, api_key, basicself serveunresolvedeasynonesource0.85
12IntercomSupport and Helpdeskoauth2, api_keyself servefirst-partyeasynonesource0.85
13FreshdeskSupport and Helpdeskapi_key, basicself servenoneeasynonesource0.75
14FrontSupport and Helpdeskapi_key, oauth2self servenoneeasynonesource0.50
15PylonSupport and Helpdeskapi_keygatedfirst-partymediumgetting a workspace to test against is sales-led (book-a-demo), even though the API itself is documentedsource0.70
16LiveAgentSupport and Helpdeskapi_keyself servenoneeasynonesource0.50
17PlainSupport and Helpdeskapi_keyself servefirst-partyeasyGraphQL-only surface (no REST) is a minor adaptation cost for a REST-oriented connector frameworksource0.90
18Help ScoutSupport and Helpdeskoauth2self servenoneeasynonesource0.50
19GorgiasSupport and Helpdeskapi_key, basic, oauth2self servenoneeasypublic App Store listing requires review (private integrations don't)source0.80
20GladlySupport and Helpdeskapi_key, basicgatednonemediumno self-serve trial — need a Gladly customer/sandbox org to obtain API tokenssource0.65
21SlackCommunications and Messagingoauth2self serveunresolvedeasyMarketplace listing requires review; granular OAuth scopes need care but plain workspace apps are trivialsource0.90
22TwilioCommunications and Messagingapi_key, basicself serveunresolvedeasyUS SMS traffic needs A2P 10DLC/toll-free registration before production sendingsource0.80
23Zoho CliqCommunications and Messagingoauth2self servenonemediumZoho OAuth + multi-datacenter endpoints; smaller API surface than Slack-class productssource0.50
24LarkCommunications and Messagingoauth2self serveunresolvedmediumapp installation/approval flow inside tenant orgs, and Lark (intl) vs Feishu (CN) platform splitsource0.50
25PumbleCommunications and Messagingapi_keyself servenonemediumlimited public API surface and small integration ecosystemsource0.50
26DiscordCommunications and Messagingoauth2, api_keyself servefirst-partyeasyprivileged intents (message content, members) require verification once a bot passes 100 serverssource0.85
27TelegramCommunications and Messagingapi_keyself servecommunityeasynonesource0.85
28WhatsApp BusinessCommunications and Messagingoauth2partialnonemediumdev access is self-serve (free Meta app, test number, access tokens) but production requires Meta business verification + message-template reviewsource0.85
29AircallCommunications and Messagingbasic, oauth2self servenoneeasyneeds a paid/trial Aircall account with phone numbers to test real call flowssource0.65
30VonageCommunications and Messagingapi_key, jwtself servenoneeasynonesource0.85
31Google AdsMarketing Ads Email and Socialoauth2partialnoneeasyExplorer Access (launched Feb 2026) is auto-granted with no formal review and allows production calls up to 2,880 operations/day; Basic (~2 day review) and Standard (~10 day review) access remain gated. API is free at all levelssource0.90
32Meta AdsMarketing Ads Email and Socialoauth2gatednonemediumMeta app review / business verification for advanced access and ads_management permissionssource0.50
33LinkedIn AdsMarketing Ads Email and Socialoauth2gatednonehardLinkedIn Marketing API Program application/approval — no self-serve production credentialssource0.90
34GoHighLevelMarketing Ads Email and Socialoauth2, api_keyself serveunresolvedmediumdocs page is JS-rendered (unfetchable here); marketplace app approval needed for public distributionsource0.50
35MailchimpMarketing Ads Email and Socialapi_key, oauth2self servenoneeasynonesource0.85
36KlaviyoMarketing Ads Email and Socialapi_key, oauth2self servefirst-partyeasynonesource0.85
37systeme.ioMarketing Ads Email and Socialapi_keyself servenoneeasynonesource0.85
38PinterestMarketing Ads Email and Socialoauth2gatednonemediumtrial access is limited; standard access application needed for production scopessource0.75
39ThreadsMarketing Ads Email and Socialoauth2gatednonemediumMeta app review for publishing permissions; limited API surfacesource0.50
40SendGridMarketing Ads Email and Socialapi_keyself servenoneeasysender identity/domain verification before sending real mailsource0.65
41ShopifyEcommerceoauth2, api_keyself servefirst-partyeasypublic App Store apps require review; GraphQL-first surface needs adaptation if your framework is REST-orientedsource0.95
42WooCommerceEcommerceapi_key, basicself servenoneeasyevery user must have their own WordPress+WooCommerce instance; no central cloud APIsource0.85
43BigCommerceEcommerceapi_key, oauth2self servenoneeasynonesource0.90
44Salesforce Commerce CloudEcommerceoauth2gatednonehardno self-serve access — requires an SFCC customer realm or partner sandbox to obtain credentialssource0.50
45MagentoEcommerceoauth2, api_key, jwtself servenonemediumper-instance APIs (self-hosted) and Adobe Commerce cloud licensing for the commercial editionsource0.50
46SquarespaceEcommerceapi_key, oauth2self servenoneeasyCommerce APIs require a paid Commerce-plan site; surface is commerce-focused, not full site managementsource0.80
47EcwidEcommerceoauth2, api_keyself servenoneeasynonesource0.85
48GumroadEcommerceoauth2, api_keyself servenonemediumnarrow, largely static API surface; docs page is JS-rendered (unfetchable here)source0.50
49Amazon Selling PartnerEcommerceoauth2gatednonehardSP-API developer registration: identity verification + role approval via Solution Provider Portal before production accesssource0.90
50fanbasisEcommercenone_publicunknownnoneblockedno public API found — integration would require a partnership conversationsource0.85
51DataForSEOData SEO and Scrapingbasicself servefirst-partyeasynonesource0.90
52SE RankingData SEO and Scrapingapi_keyself servenonemediumAPI access requires a paid Business-tier subscription; docs page unfetchable (likely bot protection)source0.50
53AhrefsData SEO and Scrapingapi_keygatedfirst-partymediumfull API access is Enterprise-plan only (free test queries available without Enterprise) — expensive gate for a general integrationsource0.90
54MrScraperData SEO and Scrapingapi_keyself servefirst-partyeasynonesource0.85
55ApifyData SEO and Scrapingapi_keyself servefirst-partyeasynonesource0.95
56FirecrawlData SEO and Scrapingapi_keyself servefirst-partyeasynonesource0.90
57Bright DataData SEO and Scrapingapi_keyself servefirst-partyeasyKYC checks apply to some proxy products; per-product pricing complexitysource0.75
58SherlockData SEO and Scrapingnone_publicself servenonemediumno hosted API — requires running the CLI yourself; results depend on scraping sites that may blocksource0.90
59Waterfall.ioData SEO and Scrapingapi_keyself servenoneeasycredit-based pricing for enrichment datasource0.80
60ClayData SEO and Scrapingapi_keygatedfirst-partymediumAPI access is tied to paid Clay workspaces; public standalone API surface is limitedsource0.50
61GitHubDeveloper Infra and Data platformsoauth2, api_key, jwtself servefirst-partyeasynonesource0.95
62VercelDeveloper Infra and Data platformsapi_key, oauth2self servefirst-partyeasynonesource0.90
63NetlifyDeveloper Infra and Data platformsoauth2, api_keyself servefirst-partyeasynonesource0.90
64CloudflareDeveloper Infra and Data platformsapi_keyself servefirst-partyeasynonesource0.95
65SupabaseDeveloper Infra and Data platformsapi_key, oauth2, jwtself servefirst-partyeasydata APIs are per-project (each user brings their own project URL + keys)source0.90
66Neo4jDeveloper Infra and Data platformsoauth2, basicself serveunresolvedmediumdatabase-centric access model (drivers/Cypher) rather than a typical SaaS REST surfacesource0.50
67SnowflakeDeveloper Infra and Data platformsoauth2, jwt, api_keyself serveunresolvedmediumaccount-specific URLs, warehouse/role setup, and auth complexity make a generic connector non-trivialsource0.70
68MongoDB AtlasDeveloper Infra and Data platformsoauth2, api_keyself servefirst-partyeasydocument CRUD requires driver connections, not REST — admin operations are the natural API integrationsource0.80
69DatadogDeveloper Infra and Data platformsapi_keyself servefirst-partyeasymulti-site endpoints (US/EU/gov) need a site parametersource0.90
70SentryDeveloper Infra and Data platformsapi_key, oauth2self servefirst-partyeasynonesource0.90
71NotionProductivity and Project Managementoauth2, api_keyself servefirst-partyeasynonesource0.95
72AirtableProductivity and Project Managementoauth2, api_keyself servefirst-partyeasynonesource0.85
73LinearProductivity and Project Managementoauth2, api_keyself servefirst-partyeasyGraphQL-only surface if your connector framework is REST-orientedsource0.90
74JiraProductivity and Project Managementoauth2, api_key, basic, jwtself serveunresolvedeasyMarketplace apps require review; granular OAuth scopes migrationsource0.90
75AsanaProductivity and Project Managementoauth2, api_keyself serveunresolvedeasynonesource0.90
76Monday.comProductivity and Project Managementapi_key, oauth2self servefirst-partyeasyGraphQL-only surface; complexity-point rate limitingsource0.90
77ClickUpProductivity and Project Managementoauth2, api_keyself servefirst-partyeasynonesource0.90
78CodaProductivity and Project Managementapi_keyself servenoneeasyrebrand already live in docs (Coda → Superhuman Docs API) — expect naming/domain churn and possible endpoint migrationsource0.95
79SmartsheetProductivity and Project Managementapi_key, oauth2self servefirst-partyeasynonesource0.90
80HarvestProductivity and Project Managementapi_key, oauth2self servenoneeasynonesource0.85
81StripeFinance and Fintechapi_keyself servefirst-partyeasynonesource0.95
82PlaidFinance and Fintechapi_keygatedfirst-partymediumproduction access requires Plaid approval and compliance review — sandbox keys are instant, live bank data is notsource0.95
83BinanceFinance and Fintechapi_keyself servenonemediumKYC-verified account required; regional restrictions (Binance.US vs global) and request-signing complexitysource0.90
84Paygent ConnectFinance and Fintechapi_keygatednonehardmerchant contract required before any credentials or full documentationsource0.50
85iPayXFinance and Fintechnone_publicunknownnoneblockedno verifiable public API documentation found — obscure product, needs direct vendor contactsource0.50
86QuickBooksFinance and Fintechoauth2self servenonemediumproduction access runs through the Intuit App Partner Program with tiered pricing: writes free, reads metered, free Builder tier at 500K reads/month; dev account + sandbox are instantsource0.90
87XeroFinance and Fintechoauth2self servefirst-partyeasyuncertified apps are capped at 25 connections — App Partner certification needed to scalesource0.85
88BrexFinance and Fintechapi_key, oauth2self servefirst-partyeasyAPI tokens require being a Brex customer; partner OAuth apps need Brex approvalsource0.80
89RampFinance and Fintechoauth2self servenoneeasyrequires a Ramp customer account to create API clientssource0.70
90PitchBookFinance and Fintechapi_keygatednonehardenterprise sales contract and data licensing required for any API accesssource0.50
91NotebookLMAI Research and Mediaoauth2gatednonehardno consumer/self-serve path — Enterprise API only, behind Google Cloud (Agentspace) contractssource0.90
92Otter AIAI Research and Medianone_publicgatednonehardno public developer program; would rely on unofficial/reverse-engineered endpoints or a partnershipsource0.50
93FathomAI Research and Mediaapi_keyself servefirst-partyeasyAPI is relatively new; surface is read-oriented (recordings, transcripts, summaries)source0.60
94ConsensusAI Research and Mediaapi_keyunknownfirst-partymediumhow API keys are issued (self-serve vs contact) isn't stated on the reference page — verify pricing/accesssource0.60
95ReductoAI Research and Mediaapi_keyself servefirst-partyeasycredit-based billing; enterprise features (on-prem, VPC) are sales-led but the API itself is self-servesource0.90
96DevinAI Research and Mediaapi_keyself servefirst-partyeasypaid subscription required; API drives an agent product rather than exposing granular data resourcessource0.75
97higgsfieldAI Research and Mediaapi_keyself servenoneeasyasync job lifecycle (submit → poll/webhook) needs slightly more plumbing than sync RESTsource0.85
98Mermaid CLIAI Research and Medianone_publicself servenonemediumno hosted API — integration means bundling the CLI + Chromium in your own execution environmentsource0.50
99YouTube TranscriptAI Research and Mediaapi_keyself servefirst-partyeasythird-party scraper of YouTube — ToS/stability risk sits with the vendorsource0.85
100GrainAI Research and Mediaapi_key, oauth2self servenoneeasynonesource0.90

3  The agent

what was built, and who did which part

A two-stage pipeline: a pure-Python fetch layer that never calls a model, and an extraction stage run by Claude. Nothing about the assessment is hidden behind an API key — the fetch stage is reproducible with pip install requests.

  1. 1

    apps.csvpython

    100 apps with id, name, category, hint_url. Many hint URLs are homepages, not developer docs — which the next step has to survive.
    → 100 rows
  2. 2

    Fetch + docs-likeness scoringpython

    Fetch the hint URL, strip HTML with the stdlib parser, then score the text against a keyword set (api reference, oauth, endpoint, rate limit, webhook…). Two or more hits = docs-like. If the page fails or reads as marketing, probe developers.X, developer.X, docs.X, X/api, X/developers and keep the first docs-like hit. 10 parallel workers, resumable via fetch_log.json.
    → 63 docs-like · 29 non-docs · 8 failed  |  pages/*.txt (≤30K chars)
  3. 3

    Extraction, in batches of 10claude

    The extraction engine is Claude, running as Claude Code and orchestrated in batches of ten pages. Each app is scored on a fixed schema — auth methods, access model, API surface, MCP, verdict, blocker, evidence URL, confidence — and merged into results.json through a small idempotent merge script.
    → results.json (100 records, one JSON object per app)
  4. 4

    Confidence cappingclaude

    The rule that made the pipeline auditable: a docs-grounded page allows confidence 0.60–0.95, but a marketing page, JS-rendered shell, or failed fetch caps confidence at 0.50 — no exceptions, even when the answer felt obvious. Verified later: capped rows erred at 5× the rate of docs-grounded rows, so the signal was real.
    → 31 rows capped at ≤0.50
  5. 5

    needs_human routingclaude

    Every capped or failed row is written to needs_human.json with the reason (JS-rendered, bot-blocked, marketing-only). This is the queue a reviewer works, rather than re-reading all 100.
    → needs_human.json
  6. 6

    Verification loopshumanclaude

    A 20-app stratified sample — 10 low-confidence vs 10 high-confidence — checked claim-by-claim across three independent methods. Corrections flow back into results.json; results_firstpass.json is frozen beforehand so the before/after is provable.
    → 80 claims checked · 6 corrections · 2 schema changes

Where a human was genuinely needed

  • Meta bot-blocking. WhatsApp, Meta Ads and Threads all return errors to an agent fetch but open fine in a human browser. Three apps only a person could confirm.
  • JS-rendered docs. GoHighLevel (Stoplight), Gumroad, QuickBooks, Lark and Ramp served shells of 3–62 characters. Ironically Ramp's shell tells agents to fetch llms.txt instead — a pattern more vendors should copy.
  • Homepage-only hint URLs. Salesforce, HubSpot and Zoho gave marketing pages; the fallback prober missed because their docs live behind different paths.
  • Judgment on schema shape. Two fields were the wrong shape, and no amount of fetching would have revealed it — only a human noticing that “dev self-serve, production gated” had nowhere to go.

Design decisions worth defending

  • Fetch and reason as separate stages. Pages are cached to disk, so re-assessment never re-fetches and the evidence is inspectable after the fact.
  • Resumability everywhere. Both stages skip completed work; an interrupted run costs nothing.
  • Confidence is self-reported and then tested. A confidence number nobody checks is decoration. The verification sample exists to price it.
  • Absence of evidence recorded as absence. Where a page was silent, the answer is unresolved — not a guess. 12 MCP origins sit in that state.

4  Verification

20 apps · 80 claims · 3 methods

A stratified sample, not a convenience sample: 10 apps the pipeline flagged as low-confidence (Group A) against 10 it was confident about (Group B), with all four core claims checked per app. The question being tested is whether the pipeline's own confidence score predicts its error rate.

Claim accuracy by group
GroupClaimsCorrectWrongAccuracyRead
A — flagged, memory-based40355 87.5%capped at 0.50 by the pipeline itself
B — high-confidence, docs-grounded40391 97.5%grounded in a fetched docs page
Total80746 92.5%Group A erred at Group B's rate

The confidence signal worked. The pipeline flagged its own weak rows without human input, and those rows were where the errors actually were. But low confidence did not mean wrong — 35 of 40 Group A claims were correct. Most flags were evidence gaps, not bad judgement: Binance and Sherlock both scored 4/4 from a failed fetch.

Per-claim accuracy

ClaimAccuracy
mcp_exists200100%
auth_methods19195.0%
self_serve_or_gated18290.0%
main_blocker17385.0%

Structured enum fields were near-perfect. The free-text field was the weakest — it absorbs both staleness and unsupported detail.

Before → after

Group AGroup B
Mean confidence before0.500.86
Mean confidence after0.890.93
Apps with ≥1 correction4 of 101 of 10

results_firstpass.json is the frozen pre-correction snapshot, so every delta on this page is reproducible from a diff.

All six errors, stated openly

Four of six were staleness, and every staleness error ran the same direction: the platform had become more open than the model remembered. Memory-based assessment systematically under-rates buildability — which is the more useful thing to know than the headline accuracy number.

AppClaimTypeCorrectionWhy it was wrong
Google Adsself_serve_or_gatedstalenessgatedpartialExplorer Access (Feb 2026) is auto-granted with no review, allowing production calls up to 2,880 ops/day.
Google Adsmain_blockerstalenessreview requirement removedOnly Basic (~2 day) and Standard (~10 day) tiers are gated. The API is free at all levels.
Salesforcemain_blockerstalenessConnected Apps → External Client AppsRenamed in Spring '26. Terminology only — the OAuth-app-setup blocker itself still stands.
NotebookLMauth_methodsstalenessnone_publicoauth2The NotebookLM Enterprise API shipped Sept 2025 via Google Cloud, so “no public auth” was wrong. Verdict also moved blockedhard.
WhatsApp Businessself_serve_or_gatedschema limitationgatedpartialDev access is self-serve (free Meta app, test number, tokens); only production is gated. The binary field could not express this.
Amazon SP-APImain_blockerunsupported specificsuncitable details stripped“Can take weeks”, “professional seller account” and “security questionnaire” were not on the page — memory-derived embellishment inside a docs-grounded row.

Two further errors fell outside the four-claim protocol, both on buildability_verdict: NotebookLM blockedhard, and Google Ads mediumeasy. Not one of the eight was a reasoning failure about a correctly-read page.

Three independent verification layers

  • Human browser (4 apps) — Salesforce, HubSpot, WhatsApp Business, Jira. Required where an agent physically cannot reach the page (Meta) or the docs are JS-rendered.
  • Second-agent pass, independent web search (6 apps) — Google Ads, QuickBooks, Binance, NotebookLM, Sherlock, fanbasis. A different engine from the one that produced the assessment, so it does not inherit the same stale priors — which is exactly how the Google Ads and NotebookLM staleness surfaced.
  • Agent re-check via live fetch (10 apps) — the whole of Group B. Every claim re-checked against a freshly fetched page; all 10 fetches succeeded, so nothing was confirmed from memory.

The honest non-results

  • 12 unresolved MCP origins. A server exists but first-party vs community could not be established — Asana, Neo4j, Snowflake, Zendesk and Twilio all 404'd on probe. Recorded as unresolved rather than guessed.
  • 33 rows still in needs_human. This rose from 31: ten sample apps closed, but the sharper MCP question opened twelve. The count went up while the data got better.
  • 21 rows still below 0.60 confidence. 79 of 100 apps have never been independently checked — the 92.5% figure is a sample estimate, not a guarantee about every row.
  • One claim was resolved only by changing the schema. Telegram's MCP status was unanswerable as a boolean; it became answerable as community.

5  What I'd fix

known limitations, in priority order

These are the changes I would make before running this over 1,000 apps instead of 100.

1. Constrain main_blocker

At 85% it is the least reliable field — free text invites both stale phrasing (Salesforce) and specifics the page never supported (Amazon SP-API). Fix: a controlled vocabulary (approval_gate, commercial, architecture, surface_risk, none) plus a mandatory verbatim quote from the source page for any free-text elaboration. If it can't be quoted, it can't be claimed.

2. Split the access field

The binary self_serve_or_gated was the wrong shape and produced 2 of the 6 errors. “Dev access self-serve, production gated” is the real pattern for Google Ads, WhatsApp, Meta Ads, Plaid, QuickBooks and Pinterest. Fix: replace with dev_access + prod_access. I added a stopgap partial value, but two fields is the correct model.

3. Schedule re-runs — this data decays

In a 100-app sample across roughly 13 months: two rebrands (FanBasis → Commas, Coda → Superhuman Docs) and two access-model changes (Google Ads Explorer Access, QuickBooks tiered pricing). A static catalogue is wrong within a quarter. Fix: treat this as a monitored dataset — re-fetch on a cadence, diff against the previous run, and alert on name/domain/access changes rather than re-reading everything.

4. Add a headless-browser tier

8 of 8 fetch failures and most “non-docs” misses were JS-rendered shells or bot-blocking, not missing documentation. Fix: a three-tier fetch — plain HTTP, then llms.txt / /openapi.json probing (already published by Ramp, Gorgias, BigCommerce, Asana, Klaviyo), then headless Chromium. That would likely close most of the 33-row human queue without a human.

5. Tighten the docs-likeness heuristic

The ≥2-keyword rule produced false negatives on real documentation — Telegram, LinkedIn, Snowflake, Harvest, Airtable and Ramp were all scored “non-docs” despite being genuine developer pages, which needlessly depressed their confidence. Fix: weight structural signals (presence of an OpenAPI link, code blocks, HTTP verbs in headings) above bare keyword counts.

6. Widen the verification sample

20 of 100 apps were verified, so the 92.5% has a wide confidence interval and Group B's single error makes its 97.5% especially soft. Fix: verify a fixed random 10% every re-run, weighted toward rows whose values changed since the last run — continuous sampling rather than one audit.