Methodology
How we measure developer adoption.
If procurement or press asks “how do you know?”, this page explains the implemented evidence, scoring, privacy, and benchmark rules behind the product.
Trust labels used across Tokens&
These are the concepts buyers, builders, and agents should see whenever a score, proof badge, ranking, report, or War Room signal appears.
Verified adoption
A usage, activation, retention, or workflow signal is verified only when it is backed by first-party telemetry, approved public proof, or a customer-supplied source that can be audited.
Read policyAccount intent score
Account intent scores summarize whether a company domain shows adoption behavior that is strong enough for GTM review.
Read policyCategory movement score
Category movement combines retained builders, territory strength, competitive pressure, white-space demand, and benchmark freshness into one operating signal.
Read policyProof freshness
Every score, badge, proof profile, report, and War Room signal should disclose whether the evidence is live, sample, mixed, stale, or insufficient.
Read policyPrivacy thresholds
Tokens& is built around aggregate and account-level proof by default. Raw developer resale is not part of the product contract.
Read policyWhere the data comes from
Two sources feed the graph. Directory-side: public signals such as tool-page views, search queries, challenge submissions, claims, and self-reported stacks. Enterprise-side: events sent by opted-in organisations through our SDK, the /ingest/events API, or marketplace connectors.
Enterprise evidence is scoped to an organisation and timestamped. An event can include a linked developer identity, but anonymous events are permitted. Source-specific receipts and verification fields determine how much weight a signal receives; the product does not treat every event as verified.
Privacy guardrails
Cross-tenant percentile benchmarks use opted-in organisations only. Publication requires at least seven distinct contributing organisations and also honours each contributor's configured minimum cohort size. Cohorts below the effective threshold are suppressed instead of extrapolated.
Cross-tenant benchmark rows contain aggregate metric values, not raw developer rows or developer identifiers. Tenant-private comparisons can use an organisation's own value without adding that value to the shared distribution.
Changing an organisation to opted out excludes it the next time the benchmark aggregation completes. That aggregation also removes a current-period cohort when the remaining contributors no longer meet the publication threshold.
Adoption index formula
The Adoption Index is a 0-100 weighted composite: adoption evidence 40%, benchmark performance 25%, review rating 20%, GitHub stars 10%, and funding/featured context 5%. The adoption component applies source-specific multipliers, caps self-reported volume, and includes bounded verification, velocity, and retention adjustments.
Category rankings, percentiles, badges and the /rankings page order tools by AgentRank, a second 0-100 composite over the Adoption Index: Adoption Index 35%, trust 25% (verified publisher, telemetry, protocol support, security evidence, freshness, reputation), ecosystem momentum 15%, benchmark performance 15%, and commercial readiness 10% (verified publisher, usage telemetry, protocols, security evidence, an enterprise or custom plan, verified adoption claims). Neither composite contains a customer-tier or funding term. Sparklines on /rankings show the last six Adoption Index captures, never a modelled line.
Scores use the latest available evidence when they are computed. The current implementation does not promise a weekly refresh cadence or apply time-decay to rankings, so period and refresh timestamps should accompany any comparison.
Public-signal basis. While a product's verified developer cohort is below the N=7 threshold, its verified metrics stay private and the public AgentRank is ordered on open-web evidence only: GitHub stars, forks and 30-day star velocity, public reviews, linked docs, API availability, and protocol support (MCP, SDK). Public-signal scores are capped at 74 so a product with a verified cohort can always rank above one that only has stars. Rows ranked this way are labelled "Public signals" on the scoreboard and carry scoreBasis="public_signals" in the API; they never disclose anything about the private cohort.
Cohorts and benchmarks
Category benchmarks (p25, p50, p75, p90) are only published when the opted-in cohort meets the effective minimum. If a requested cohort is below that threshold, the API returns a suppressed result rather than extrapolating.
Each organisation contributes one averaged value per metric to a distribution, regardless of how many products or metric rows it has. Available retention benchmarks currently include D7 and D30 rates from retention snapshots; there is no separate ten-developer-per-organisation publication rule.
What we do not claim
We do not claim causation. The product attributes developer activity to a program with a lift estimate and confidence interval. Where there is not enough signal, we surface that explicitly.
We do not rank on money. Customer tier does not influence ranking or benchmark inclusion. Ranking is a function of the data, full stop.
Audit and versioning
Methodology changes are versioned in source control and release history. We do not currently promise a numeric-change threshold, advance-notice window, or automatic re-publication SLA; a cited result should therefore include its methodology version, period, and refresh time.
Our SOC 2 Type I report is available on request under NDA. Type II is in progress.
