Dataset
165 models · 4 languages[1]
Reported by the paper for English, Chinese, Russian, and Arabic testing. This is not a ChinaAPI dataset.
Paper · data recordA published black-box method shows that repeated one-token outputs can form a measurable behavioral fingerprint. ChinaAPI is prototyping a reproducible consistency audit for API routes.
Dataset
165 models · 4 languagesReported by the paper for English, Chinese, Russian, and Arabic testing. This is not a ChinaAPI dataset.
Paper · data recordAudit cost
~100 one-token queriesThe paper reports this approximate request count for an 8-cell protocol. Actual API cost and evidence strength depend on the endpoint and configuration.
PaperVerification result
<11% EER with 8 probe cellsA result reported under the paper's experimental protocol. It is not a guaranteed error rate for arbitrary routes, prompts, or providers.
PaperWhy one response proves nothing
A model identifier tells you what a route claims to serve. It is not independent evidence of the route's behavior.
Random sampling, service updates, system prompts, and routing policy can all change one completion.
Response time can expose infrastructure conditions, but it cannot by itself identify a model.
The usable evidence comes from the shape of repeated normalized answers, not from any single answer.
What the paper actually did
The researchers used constrained prompts designed to yield one-word answers, repeated the sampling, normalized the returned answer categories, and compared the resulting probability-like fingerprints.
Ask questions with a small, controlled answer space so repeated outputs can be counted consistently.
Collect many independent one-token outputs under fixed endpoint and sampling settings.
Map equivalent surface forms into stable answer categories before comparison.
Measure whether two answer distributions are closer than expected for different models.
Independent splits keep a similar shape.
Peaks and category weights diverge.
Text alternative: the left panel shows three different illustrative bar distributions. The middle overlays two independently sampled distributions from the same illustrative model and shows similar bar heights. The right overlays two different illustrative models and shows distinct peaks and category weights. No real model measurements are shown.
Key research results
The paper found that same-model fingerprints remain substantially closer than fingerprints from different models.
When samples from the same model were split in half, their fingerprints were about an order of magnitude closer than samples from different models. This comparison is reported by the paper under its probe construction and distance metrics.[1]
The paper also reports an ecosystem anomaly in which one endpoint carrying a proprietary flagship label was distributionally indistinguishable from open-source Qwen models. This is the paper authors' reported observation, not a ChinaAPI test or attribution claim.[1]
| Paper result | Reported value |
|---|---|
Model-family classification[1]Leave-one-out accuracy reported by the paper. Paper | 59.5% |
Chance baseline[1]Random baseline reported alongside the classification result. Paper | 18.4% |
Full 40-cell verification EER[1]Equal error rate under the paper's complete 40-cell protocol. Paper | 7.3% |
8-cell verification EER[1]Reported with roughly 100 one-token requests; not a universal route guarantee. Paper | <11% |
All values are reported by One Token Is Enough; the associated research record is available on Zenodo.
Proposed private beta
The private beta will expose the evidence needed to reproduce and challenge a comparison. The report is designed to retain method context, not collapse the result into an unexplained score.
Request parameters and the exact probe-set version.
Test timestamps, route, configuration, and sample count.
Raw responses alongside their normalized categories.
The requested model name and any model identifier returned by the API.
Jensen-Shannon divergence from a versioned reference fingerprint.
Confidence interval, anomaly type, and change from prior windows.
How to interpret a report
Consistent with reference
Observed differences remain within the expected sampling range for this configuration.
Insufficient evidence
The current sample or task panel cannot support a reliable comparison.
Mismatch signal detected
Multiple independent probe cells differ from the reference beyond the stated threshold. Further investigation is warranted.
Method boundaries
Beta data boundary
The first beta is intended to evaluate requests made under a participant's own ChinaAPI account. A later Dashboard workflow may offer short-lived, restricted test credentials. This page will not ask for a long-lived key from another provider.
Interest-form details are used to evaluate and contact beta candidates. If no beta relationship begins, the inquiry is scheduled for deletion within 90 days; participants may request earlier deletion at [email protected]. If a participant joins, the beta agreement shown before testing will state the audit-record retention period and deletion workflow. See the Privacy Policy.
Private beta interest
Tell us which routes and workloads matter. Do not include an API key, prompt contents, customer data, or other secrets.
Research preview
Join the ChinaAPI private beta to help shape a reproducible standard for model-route consistency testing.