94 of 100 apps authenticate with an API key, OAuth 2.0, or both (39 support both). Two auth paths cover the entire catalogue. Auth is not where the engineering cost lives — so don't budget for it as if it were.
Which of these 100 apps can you actually integrate?
An agent-built buildability assessment of 100 third-party apps — auth model, whether a developer can self-serve credentials, whether an MCP server already exists, and what actually blocks the build. Every claim carries an evidence URL and a confidence score; a 20-app verification sample puts claim accuracy at 92.5%.
1 Headline patterns
what the 100 rows add up toEach number below is followed by what it means for someone building an integration toolkit — the interpretation is the point, not the count.
35 apps are easy + self-serve with no blocker at all — documented REST API, instant credentials, nothing to negotiate. Start here: roughly a third of the catalogue can ship without a single conversation.
Of the 65 apps with a blocker, 30 are approval gates — app review, partner programmes, sales contracts, developer registration. Only 12 are architectural. The bottleneck is a partnerships motion, not engineering time.
40 vendors ship and maintain their own MCP server; 1 is community-only. In dev-infra and productivity the surface is already built — adopt, don't rebuild. Ecommerce, comms and ads are wide open by contrast.
Self-serve rate by category. Productivity and dev-infra are 20/20 self-serve and 18/20 easy; marketing/ads is 5/10 self-serve with 4 hard gates (Google, Meta, LinkedIn, Pinterest). Sequence the roadmap by category — and staff ads platforms with BD, not just engineers.
21 apps are gated, partial, or unknown-access; 10 are hard or blocked outright. Qualify these before committing sprint capacity — 2 (fanbasis, iPayX) have no public API at all and can't be built against today.
2 The 100
filter, search, sort · ✓ = independently verifiedEvery row is one assessed app. The confidence bar is the model's own self-reported confidence at extraction time; anything under 0.60 was routed to human review.
| # ▲▼ | App ▲▼ | Category ▲▼ | Auth | Access ▲▼ | MCP ▲▼ | Verdict ▲▼ | Main blocker | Evidence | Conf. ▲▼ |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Salesforce✓ | CRM and Sales | oauth2, jwt | self serve | unresolved | medium | External Client App (formerly Connected App, renamed Spring '26) + OAuth setup and per-org instance URLs add integration complexity | source | 0.85 |
| 2 | HubSpot✓ | CRM and Sales | oauth2, api_key | self serve | first-party | easy | none | source | 0.90 |
| 3 | Pipedrive | CRM and Sales | api_key, oauth2 | self serve | unresolved | easy | none | source | 0.85 |
| 4 | Attio | CRM and Sales | oauth2, api_key | self serve | first-party | easy | none | source | 0.80 |
| 5 | Twenty | CRM and Sales | api_key | self serve | none | easy | smaller ecosystem; API surface still evolving with the open-source project | source | 0.65 |
| 6 | Podio | CRM and Sales | oauth2, api_key | self serve | none | easy | product is in low-investment/maintenance mode, so long-term API viability is a risk | source | 0.80 |
| 7 | Zoho CRM | CRM and Sales | oauth2 | self serve | unresolved | medium | multi-datacenter endpoints (zoho.com/.eu/.in) and OAuth-only auth complicate a generic connector | source | 0.50 |
| 8 | Close | CRM and Sales | api_key, oauth2 | self serve | first-party | easy | none | source | 0.70 |
| 9 | Copper | CRM and Sales | api_key, oauth2 | self serve | none | easy | none | source | 0.85 |
| 10 | DealCloud | CRM and Sales | api_key, oauth2 | gated | first-party | hard | enterprise product — API keys exist per client site, so you need a DealCloud customer instance to build/test against | source | 0.70 |
| 11 | Zendesk | Support and Helpdesk | oauth2, api_key, basic | self serve | unresolved | easy | none | source | 0.85 |
| 12 | Intercom | Support and Helpdesk | oauth2, api_key | self serve | first-party | easy | none | source | 0.85 |
| 13 | Freshdesk | Support and Helpdesk | api_key, basic | self serve | none | easy | none | source | 0.75 |
| 14 | Front | Support and Helpdesk | api_key, oauth2 | self serve | none | easy | none | source | 0.50 |
| 15 | Pylon | Support and Helpdesk | api_key | gated | first-party | medium | getting a workspace to test against is sales-led (book-a-demo), even though the API itself is documented | source | 0.70 |
| 16 | LiveAgent | Support and Helpdesk | api_key | self serve | none | easy | none | source | 0.50 |
| 17 | Plain | Support and Helpdesk | api_key | self serve | first-party | easy | GraphQL-only surface (no REST) is a minor adaptation cost for a REST-oriented connector framework | source | 0.90 |
| 18 | Help Scout | Support and Helpdesk | oauth2 | self serve | none | easy | none | source | 0.50 |
| 19 | Gorgias | Support and Helpdesk | api_key, basic, oauth2 | self serve | none | easy | public App Store listing requires review (private integrations don't) | source | 0.80 |
| 20 | Gladly | Support and Helpdesk | api_key, basic | gated | none | medium | no self-serve trial — need a Gladly customer/sandbox org to obtain API tokens | source | 0.65 |
| 21 | Slack | Communications and Messaging | oauth2 | self serve | unresolved | easy | Marketplace listing requires review; granular OAuth scopes need care but plain workspace apps are trivial | source | 0.90 |
| 22 | Twilio | Communications and Messaging | api_key, basic | self serve | unresolved | easy | US SMS traffic needs A2P 10DLC/toll-free registration before production sending | source | 0.80 |
| 23 | Zoho Cliq | Communications and Messaging | oauth2 | self serve | none | medium | Zoho OAuth + multi-datacenter endpoints; smaller API surface than Slack-class products | source | 0.50 |
| 24 | Lark | Communications and Messaging | oauth2 | self serve | unresolved | medium | app installation/approval flow inside tenant orgs, and Lark (intl) vs Feishu (CN) platform split | source | 0.50 |
| 25 | Pumble | Communications and Messaging | api_key | self serve | none | medium | limited public API surface and small integration ecosystem | source | 0.50 |
| 26 | Discord | Communications and Messaging | oauth2, api_key | self serve | first-party | easy | privileged intents (message content, members) require verification once a bot passes 100 servers | source | 0.85 |
| 27 | Telegram✓ | Communications and Messaging | api_key | self serve | community | easy | none | source | 0.85 |
| 28 | WhatsApp Business✓ | Communications and Messaging | oauth2 | partial | none | medium | dev access is self-serve (free Meta app, test number, access tokens) but production requires Meta business verification + message-template review | source | 0.85 |
| 29 | Aircall | Communications and Messaging | basic, oauth2 | self serve | none | easy | needs a paid/trial Aircall account with phone numbers to test real call flows | source | 0.65 |
| 30 | Vonage | Communications and Messaging | api_key, jwt | self serve | none | easy | none | source | 0.85 |
| 31 | Google Ads✓ | Marketing Ads Email and Social | oauth2 | partial | none | easy | Explorer Access (launched Feb 2026) is auto-granted with no formal review and allows production calls up to 2,880 operations/day; Basic (~2 day review) and Standard (~10 day review) access remain gated. API is free at all levels | source | 0.90 |
| 32 | Meta Ads | Marketing Ads Email and Social | oauth2 | gated | none | medium | Meta app review / business verification for advanced access and ads_management permissions | source | 0.50 |
| 33 | LinkedIn Ads✓ | Marketing Ads Email and Social | oauth2 | gated | none | hard | LinkedIn Marketing API Program application/approval — no self-serve production credentials | source | 0.90 |
| 34 | GoHighLevel | Marketing Ads Email and Social | oauth2, api_key | self serve | unresolved | medium | docs page is JS-rendered (unfetchable here); marketplace app approval needed for public distribution | source | 0.50 |
| 35 | Mailchimp | Marketing Ads Email and Social | api_key, oauth2 | self serve | none | easy | none | source | 0.85 |
| 36 | Klaviyo | Marketing Ads Email and Social | api_key, oauth2 | self serve | first-party | easy | none | source | 0.85 |
| 37 | systeme.io | Marketing Ads Email and Social | api_key | self serve | none | easy | none | source | 0.85 |
| 38 | Marketing Ads Email and Social | oauth2 | gated | none | medium | trial access is limited; standard access application needed for production scopes | source | 0.75 | |
| 39 | Threads | Marketing Ads Email and Social | oauth2 | gated | none | medium | Meta app review for publishing permissions; limited API surface | source | 0.50 |
| 40 | SendGrid | Marketing Ads Email and Social | api_key | self serve | none | easy | sender identity/domain verification before sending real mail | source | 0.65 |
| 41 | Shopify✓ | Ecommerce | oauth2, api_key | self serve | first-party | easy | public App Store apps require review; GraphQL-first surface needs adaptation if your framework is REST-oriented | source | 0.95 |
| 42 | WooCommerce | Ecommerce | api_key, basic | self serve | none | easy | every user must have their own WordPress+WooCommerce instance; no central cloud API | source | 0.85 |
| 43 | BigCommerce | Ecommerce | api_key, oauth2 | self serve | none | easy | none | source | 0.90 |
| 44 | Salesforce Commerce Cloud | Ecommerce | oauth2 | gated | none | hard | no self-serve access — requires an SFCC customer realm or partner sandbox to obtain credentials | source | 0.50 |
| 45 | Magento | Ecommerce | oauth2, api_key, jwt | self serve | none | medium | per-instance APIs (self-hosted) and Adobe Commerce cloud licensing for the commercial edition | source | 0.50 |
| 46 | Squarespace | Ecommerce | api_key, oauth2 | self serve | none | easy | Commerce APIs require a paid Commerce-plan site; surface is commerce-focused, not full site management | source | 0.80 |
| 47 | Ecwid | Ecommerce | oauth2, api_key | self serve | none | easy | none | source | 0.85 |
| 48 | Gumroad | Ecommerce | oauth2, api_key | self serve | none | medium | narrow, largely static API surface; docs page is JS-rendered (unfetchable here) | source | 0.50 |
| 49 | Amazon Selling Partner✓ | Ecommerce | oauth2 | gated | none | hard | SP-API developer registration: identity verification + role approval via Solution Provider Portal before production access | source | 0.90 |
| 50 | fanbasis✓ | Ecommerce | none_public | unknown | none | blocked | no public API found — integration would require a partnership conversation | source | 0.85 |
| 51 | DataForSEO | Data SEO and Scraping | basic | self serve | first-party | easy | none | source | 0.90 |
| 52 | SE Ranking | Data SEO and Scraping | api_key | self serve | none | medium | API access requires a paid Business-tier subscription; docs page unfetchable (likely bot protection) | source | 0.50 |
| 53 | Ahrefs✓ | Data SEO and Scraping | api_key | gated | first-party | medium | full API access is Enterprise-plan only (free test queries available without Enterprise) — expensive gate for a general integration | source | 0.90 |
| 54 | MrScraper | Data SEO and Scraping | api_key | self serve | first-party | easy | none | source | 0.85 |
| 55 | Apify | Data SEO and Scraping | api_key | self serve | first-party | easy | none | source | 0.95 |
| 56 | Firecrawl | Data SEO and Scraping | api_key | self serve | first-party | easy | none | source | 0.90 |
| 57 | Bright Data | Data SEO and Scraping | api_key | self serve | first-party | easy | KYC checks apply to some proxy products; per-product pricing complexity | source | 0.75 |
| 58 | Sherlock✓ | Data SEO and Scraping | none_public | self serve | none | medium | no hosted API — requires running the CLI yourself; results depend on scraping sites that may block | source | 0.90 |
| 59 | Waterfall.io | Data SEO and Scraping | api_key | self serve | none | easy | credit-based pricing for enrichment data | source | 0.80 |
| 60 | Clay | Data SEO and Scraping | api_key | gated | first-party | medium | API access is tied to paid Clay workspaces; public standalone API surface is limited | source | 0.50 |
| 61 | GitHub✓ | Developer Infra and Data platforms | oauth2, api_key, jwt | self serve | first-party | easy | none | source | 0.95 |
| 62 | Vercel | Developer Infra and Data platforms | api_key, oauth2 | self serve | first-party | easy | none | source | 0.90 |
| 63 | Netlify | Developer Infra and Data platforms | oauth2, api_key | self serve | first-party | easy | none | source | 0.90 |
| 64 | Cloudflare | Developer Infra and Data platforms | api_key | self serve | first-party | easy | none | source | 0.95 |
| 65 | Supabase | Developer Infra and Data platforms | api_key, oauth2, jwt | self serve | first-party | easy | data APIs are per-project (each user brings their own project URL + keys) | source | 0.90 |
| 66 | Neo4j | Developer Infra and Data platforms | oauth2, basic | self serve | unresolved | medium | database-centric access model (drivers/Cypher) rather than a typical SaaS REST surface | source | 0.50 |
| 67 | Snowflake | Developer Infra and Data platforms | oauth2, jwt, api_key | self serve | unresolved | medium | account-specific URLs, warehouse/role setup, and auth complexity make a generic connector non-trivial | source | 0.70 |
| 68 | MongoDB Atlas | Developer Infra and Data platforms | oauth2, api_key | self serve | first-party | easy | document CRUD requires driver connections, not REST — admin operations are the natural API integration | source | 0.80 |
| 69 | Datadog | Developer Infra and Data platforms | api_key | self serve | first-party | easy | multi-site endpoints (US/EU/gov) need a site parameter | source | 0.90 |
| 70 | Sentry | Developer Infra and Data platforms | api_key, oauth2 | self serve | first-party | easy | none | source | 0.90 |
| 71 | Notion✓ | Productivity and Project Management | oauth2, api_key | self serve | first-party | easy | none | source | 0.95 |
| 72 | Airtable | Productivity and Project Management | oauth2, api_key | self serve | first-party | easy | none | source | 0.85 |
| 73 | Linear | Productivity and Project Management | oauth2, api_key | self serve | first-party | easy | GraphQL-only surface if your connector framework is REST-oriented | source | 0.90 |
| 74 | Jira✓ | Productivity and Project Management | oauth2, api_key, basic, jwt | self serve | unresolved | easy | Marketplace apps require review; granular OAuth scopes migration | source | 0.90 |
| 75 | Asana | Productivity and Project Management | oauth2, api_key | self serve | unresolved | easy | none | source | 0.90 |
| 76 | Monday.com | Productivity and Project Management | api_key, oauth2 | self serve | first-party | easy | GraphQL-only surface; complexity-point rate limiting | source | 0.90 |
| 77 | ClickUp | Productivity and Project Management | oauth2, api_key | self serve | first-party | easy | none | source | 0.90 |
| 78 | Coda✓ | Productivity and Project Management | api_key | self serve | none | easy | rebrand already live in docs (Coda → Superhuman Docs API) — expect naming/domain churn and possible endpoint migration | source | 0.95 |
| 79 | Smartsheet | Productivity and Project Management | api_key, oauth2 | self serve | first-party | easy | none | source | 0.90 |
| 80 | Harvest | Productivity and Project Management | api_key, oauth2 | self serve | none | easy | none | source | 0.85 |
| 81 | Stripe✓ | Finance and Fintech | api_key | self serve | first-party | easy | none | source | 0.95 |
| 82 | Plaid✓ | Finance and Fintech | api_key | gated | first-party | medium | production access requires Plaid approval and compliance review — sandbox keys are instant, live bank data is not | source | 0.95 |
| 83 | Binance✓ | Finance and Fintech | api_key | self serve | none | medium | KYC-verified account required; regional restrictions (Binance.US vs global) and request-signing complexity | source | 0.90 |
| 84 | Paygent Connect | Finance and Fintech | api_key | gated | none | hard | merchant contract required before any credentials or full documentation | source | 0.50 |
| 85 | iPayX | Finance and Fintech | none_public | unknown | none | blocked | no verifiable public API documentation found — obscure product, needs direct vendor contact | source | 0.50 |
| 86 | QuickBooks✓ | Finance and Fintech | oauth2 | self serve | none | medium | production access runs through the Intuit App Partner Program with tiered pricing: writes free, reads metered, free Builder tier at 500K reads/month; dev account + sandbox are instant | source | 0.90 |
| 87 | Xero | Finance and Fintech | oauth2 | self serve | first-party | easy | uncertified apps are capped at 25 connections — App Partner certification needed to scale | source | 0.85 |
| 88 | Brex | Finance and Fintech | api_key, oauth2 | self serve | first-party | easy | API tokens require being a Brex customer; partner OAuth apps need Brex approval | source | 0.80 |
| 89 | Ramp | Finance and Fintech | oauth2 | self serve | none | easy | requires a Ramp customer account to create API clients | source | 0.70 |
| 90 | PitchBook | Finance and Fintech | api_key | gated | none | hard | enterprise sales contract and data licensing required for any API access | source | 0.50 |
| 91 | NotebookLM✓ | AI Research and Media | oauth2 | gated | none | hard | no consumer/self-serve path — Enterprise API only, behind Google Cloud (Agentspace) contracts | source | 0.90 |
| 92 | Otter AI | AI Research and Media | none_public | gated | none | hard | no public developer program; would rely on unofficial/reverse-engineered endpoints or a partnership | source | 0.50 |
| 93 | Fathom | AI Research and Media | api_key | self serve | first-party | easy | API is relatively new; surface is read-oriented (recordings, transcripts, summaries) | source | 0.60 |
| 94 | Consensus | AI Research and Media | api_key | unknown | first-party | medium | how API keys are issued (self-serve vs contact) isn't stated on the reference page — verify pricing/access | source | 0.60 |
| 95 | Reducto | AI Research and Media | api_key | self serve | first-party | easy | credit-based billing; enterprise features (on-prem, VPC) are sales-led but the API itself is self-serve | source | 0.90 |
| 96 | Devin | AI Research and Media | api_key | self serve | first-party | easy | paid subscription required; API drives an agent product rather than exposing granular data resources | source | 0.75 |
| 97 | higgsfield | AI Research and Media | api_key | self serve | none | easy | async job lifecycle (submit → poll/webhook) needs slightly more plumbing than sync REST | source | 0.85 |
| 98 | Mermaid CLI | AI Research and Media | none_public | self serve | none | medium | no hosted API — integration means bundling the CLI + Chromium in your own execution environment | source | 0.50 |
| 99 | YouTube Transcript | AI Research and Media | api_key | self serve | first-party | easy | third-party scraper of YouTube — ToS/stability risk sits with the vendor | source | 0.85 |
| 100 | Grain | AI Research and Media | api_key, oauth2 | self serve | none | easy | none | source | 0.90 |
3 The agent
what was built, and who did which partA two-stage pipeline: a pure-Python fetch layer that never calls a model, and an
extraction stage run by Claude. Nothing about the assessment is hidden behind an API key —
the fetch stage is reproducible with pip install requests.
- 1
apps.csvpython
100 apps withid, name, category, hint_url. Many hint URLs are homepages, not developer docs — which the next step has to survive.→ 100 rows - 2
Fetch + docs-likeness scoringpython
Fetch the hint URL, strip HTML with the stdlib parser, then score the text against a keyword set (api reference, oauth, endpoint, rate limit, webhook…). Two or more hits = docs-like. If the page fails or reads as marketing, probedevelopers.X,developer.X,docs.X,X/api,X/developersand keep the first docs-like hit. 10 parallel workers, resumable viafetch_log.json.→ 63 docs-like · 29 non-docs · 8 failed | pages/*.txt (≤30K chars) - 3
Extraction, in batches of 10claude
The extraction engine is Claude, running as Claude Code and orchestrated in batches of ten pages. Each app is scored on a fixed schema — auth methods, access model, API surface, MCP, verdict, blocker, evidence URL, confidence — and merged intoresults.jsonthrough a small idempotent merge script.→ results.json (100 records, one JSON object per app) - 4
Confidence cappingclaude
The rule that made the pipeline auditable: a docs-grounded page allows confidence 0.60–0.95, but a marketing page, JS-rendered shell, or failed fetch caps confidence at 0.50 — no exceptions, even when the answer felt obvious. Verified later: capped rows erred at 5× the rate of docs-grounded rows, so the signal was real.→ 31 rows capped at ≤0.50 - 5
needs_human routingclaude
Every capped or failed row is written toneeds_human.jsonwith the reason (JS-rendered, bot-blocked, marketing-only). This is the queue a reviewer works, rather than re-reading all 100.→ needs_human.json - 6
Verification loopshumanclaude
A 20-app stratified sample — 10 low-confidence vs 10 high-confidence — checked claim-by-claim across three independent methods. Corrections flow back intoresults.json;results_firstpass.jsonis frozen beforehand so the before/after is provable.→ 80 claims checked · 6 corrections · 2 schema changes
Where a human was genuinely needed
- Meta bot-blocking. WhatsApp, Meta Ads and Threads all return errors to an agent fetch but open fine in a human browser. Three apps only a person could confirm.
- JS-rendered docs. GoHighLevel (Stoplight), Gumroad, QuickBooks, Lark and Ramp
served shells of 3–62 characters. Ironically Ramp's shell tells agents to fetch
llms.txtinstead — a pattern more vendors should copy. - Homepage-only hint URLs. Salesforce, HubSpot and Zoho gave marketing pages; the fallback prober missed because their docs live behind different paths.
- Judgment on schema shape. Two fields were the wrong shape, and no amount of fetching would have revealed it — only a human noticing that “dev self-serve, production gated” had nowhere to go.
Design decisions worth defending
- Fetch and reason as separate stages. Pages are cached to disk, so re-assessment never re-fetches and the evidence is inspectable after the fact.
- Resumability everywhere. Both stages skip completed work; an interrupted run costs nothing.
- Confidence is self-reported and then tested. A confidence number nobody checks is decoration. The verification sample exists to price it.
- Absence of evidence recorded as absence. Where a page was silent, the answer is
unresolved— not a guess. 12 MCP origins sit in that state.
4 Verification
20 apps · 80 claims · 3 methodsA stratified sample, not a convenience sample: 10 apps the pipeline flagged as low-confidence (Group A) against 10 it was confident about (Group B), with all four core claims checked per app. The question being tested is whether the pipeline's own confidence score predicts its error rate.
| Group | Claims | Correct | Wrong | Accuracy | Read |
|---|---|---|---|---|---|
| A — flagged, memory-based | 40 | 35 | 5 | 87.5% | capped at 0.50 by the pipeline itself |
| B — high-confidence, docs-grounded | 40 | 39 | 1 | 97.5% | grounded in a fetched docs page |
| Total | 80 | 74 | 6 | 92.5% | Group A erred at 5× Group B's rate |
The confidence signal worked. The pipeline flagged its own weak rows without human input, and those rows were where the errors actually were. But low confidence did not mean wrong — 35 of 40 Group A claims were correct. Most flags were evidence gaps, not bad judgement: Binance and Sherlock both scored 4/4 from a failed fetch.
Per-claim accuracy
| Claim | ✓ | ✗ | Accuracy |
|---|---|---|---|
mcp_exists | 20 | 0 | 100% |
auth_methods | 19 | 1 | 95.0% |
self_serve_or_gated | 18 | 2 | 90.0% |
main_blocker | 17 | 3 | 85.0% |
Structured enum fields were near-perfect. The free-text field was the weakest — it absorbs both staleness and unsupported detail.
Before → after
| Group A | Group B | |
|---|---|---|
| Mean confidence before | 0.50 | 0.86 |
| Mean confidence after | 0.89 | 0.93 |
| Apps with ≥1 correction | 4 of 10 | 1 of 10 |
results_firstpass.json is the frozen pre-correction snapshot,
so every delta on this page is reproducible from a diff.
All six errors, stated openly
Four of six were staleness, and every staleness error ran the same direction: the platform had become more open than the model remembered. Memory-based assessment systematically under-rates buildability — which is the more useful thing to know than the headline accuracy number.
| App | Claim | Type | Correction | Why it was wrong |
|---|---|---|---|---|
| Google Ads | self_serve_or_gated | staleness | gated → partial | Explorer Access (Feb 2026) is auto-granted with no review, allowing production calls up to 2,880 ops/day. |
| Google Ads | main_blocker | staleness | review requirement removed | Only Basic (~2 day) and Standard (~10 day) tiers are gated. The API is free at all levels. |
| Salesforce | main_blocker | staleness | Connected Apps → External Client Apps | Renamed in Spring '26. Terminology only — the OAuth-app-setup blocker itself still stands. |
| NotebookLM | auth_methods | staleness | none_public → oauth2 | The NotebookLM Enterprise API shipped Sept 2025 via Google Cloud, so “no public auth” was wrong. Verdict also moved blocked → hard. |
| WhatsApp Business | self_serve_or_gated | schema limitation | gated → partial | Dev access is self-serve (free Meta app, test number, tokens); only production is gated. The binary field could not express this. |
| Amazon SP-API | main_blocker | unsupported specifics | uncitable details stripped | “Can take weeks”, “professional seller account” and “security questionnaire” were not on the page — memory-derived embellishment inside a docs-grounded row. |
Two further errors fell outside the four-claim protocol, both on
buildability_verdict: NotebookLM blocked → hard, and
Google Ads medium → easy. Not one of the eight was a reasoning failure
about a correctly-read page.
Three independent verification layers
- Human browser (4 apps) — Salesforce, HubSpot, WhatsApp Business, Jira. Required where an agent physically cannot reach the page (Meta) or the docs are JS-rendered.
- Second-agent pass, independent web search (6 apps) — Google Ads, QuickBooks, Binance, NotebookLM, Sherlock, fanbasis. A different engine from the one that produced the assessment, so it does not inherit the same stale priors — which is exactly how the Google Ads and NotebookLM staleness surfaced.
- Agent re-check via live fetch (10 apps) — the whole of Group B. Every claim re-checked against a freshly fetched page; all 10 fetches succeeded, so nothing was confirmed from memory.
The honest non-results
- 12 unresolved MCP origins. A server exists but first-party vs community could not
be established — Asana, Neo4j, Snowflake, Zendesk and Twilio all 404'd on probe. Recorded as
unresolvedrather than guessed. - 33 rows still in needs_human. This rose from 31: ten sample apps closed, but the sharper MCP question opened twelve. The count went up while the data got better.
- 21 rows still below 0.60 confidence. 79 of 100 apps have never been independently checked — the 92.5% figure is a sample estimate, not a guarantee about every row.
- One claim was resolved only by changing the schema. Telegram's MCP status was
unanswerable as a boolean; it became answerable as
community.
5 What I'd fix
known limitations, in priority orderThese are the changes I would make before running this over 1,000 apps instead of 100.
1. Constrain main_blocker
At 85% it is the least reliable
field — free text invites both stale phrasing (Salesforce) and specifics the page never
supported (Amazon SP-API). Fix: a controlled vocabulary
(approval_gate, commercial, architecture,
surface_risk, none) plus a mandatory verbatim quote from the source
page for any free-text elaboration. If it can't be quoted, it can't be claimed.
2. Split the access field
The binary
self_serve_or_gated was the wrong shape and produced 2 of the 6 errors.
“Dev access self-serve, production gated” is the real pattern for Google Ads,
WhatsApp, Meta Ads, Plaid, QuickBooks and Pinterest. Fix: replace with
dev_access + prod_access. I added a stopgap
partial value, but two fields is the correct model.
3. Schedule re-runs — this data decays
In a 100-app sample across roughly 13 months: two rebrands (FanBasis → Commas, Coda → Superhuman Docs) and two access-model changes (Google Ads Explorer Access, QuickBooks tiered pricing). A static catalogue is wrong within a quarter. Fix: treat this as a monitored dataset — re-fetch on a cadence, diff against the previous run, and alert on name/domain/access changes rather than re-reading everything.
4. Add a headless-browser tier
8 of 8 fetch failures and most
“non-docs” misses were JS-rendered shells or bot-blocking, not missing
documentation. Fix: a three-tier fetch — plain HTTP, then llms.txt /
/openapi.json probing (already published by Ramp, Gorgias, BigCommerce, Asana,
Klaviyo), then headless Chromium. That would likely close most of the 33-row human queue
without a human.
5. Tighten the docs-likeness heuristic
The ≥2-keyword rule produced false negatives on real documentation — Telegram, LinkedIn, Snowflake, Harvest, Airtable and Ramp were all scored “non-docs” despite being genuine developer pages, which needlessly depressed their confidence. Fix: weight structural signals (presence of an OpenAPI link, code blocks, HTTP verbs in headings) above bare keyword counts.
6. Widen the verification sample
20 of 100 apps were verified, so the 92.5% has a wide confidence interval and Group B's single error makes its 97.5% especially soft. Fix: verify a fixed random 10% every re-run, weighted toward rows whose values changed since the last run — continuous sampling rather than one audit.