How AfterQuery Helped NVIDIA Hill-Climb GDPval
A data vendor's account of NVIDIA using its tasks to lift Nemotron 3 Ultra on GDPval. What the claims are, and what they rest on.
How AfterQuery Helped NVIDIA Hill-Climb GDPval — AfterQuery, July 2, 2026.
Claim ledger
Assessments are the model’s knowledge, not verification.
- 1¶
AfterQuery claims its off-the-shelf dataset was used by NVIDIA, publicly and by name, to raise Nemotron 3 Ultra's GDPval performance.
“NVIDIA publicly used AfterQuery's Off-The-Shelf Office Agent Training Dataset to improve Nemotron 3 Ultra on GDPval.”
assertion · unclear
unverifiable · low confidence — I have no knowledge of a model called 'Nemotron 3 Ultra' or of this NVIDIA-AfterQuery partnership; this is dated mid-2026, past my training cutoff, and I cannot confirm NVIDIA's technical report language or authenticity independently. This is also the vendor's own framing of a customer relationship — a book-talking claim by nature.
To check: The NVIDIA Nemotron 3 Ultra technical report itself, and whether it names AfterQuery as stated.
- 2¶
Nemotron 3 Ultra is described as a 550B-parameter, 55B-active LatentMoE model released with open weights, data and recipes.
“Nemotron 3 Ultra is a fully open 550B-A55B LatentMoE model with open weights, training data, and recipes.”
quantity · unclear
unverifiable · low confidence — I have no record of an NVIDIA model called 'Nemotron 3 Ultra' or a '550B-A55B LatentMoE' architecture; NVIDIA's known Nemotron line (e.g., Nemotron-4, Nemotron-H, Llama-3.1-Nemotron) predates this. This falls after my knowledge cutoff so I cannot confirm or deny the spec.
To check: NVIDIA's official model card/technical report for parameter count, architecture type, and open-source release status.
- 3¶
Ultra is claimed to hit roughly 6x the throughput of comparable open models at equal accuracy on long-horizon agentic tasks, with a 1M-token context.
“Ultra runs at up to ~6× the throughput of comparable open models (5.9× vs GLM-5.1, 4.8× vs Kimi K2.6) on long-horizon agentic tasks at the same accuracy, and supports a context length of up to 1M tokens.”
quantity · unclear
unverifiable · low confidence — GLM-5.1 and Kimi K2.6 are model versions I have no record of; these appear to be releases after my training cutoff. The throughput comparison numbers cannot be checked against any benchmark I know.
To check: Independent throughput benchmarks (tokens/sec at matched accuracy) comparing Nemotron 3 Ultra, GLM-5.1, and Kimi K2.6.
- 4¶
AfterQuery claims exclusivity of naming among data partners in NVIDIA's technical report.
“AfterQuery is the only data partner named in the Nemotron 3 Ultra technical report:”
assertion · unclear
unverifiable · low confidence — This is a negative/exclusivity claim ('only') that the post supports merely by quoting one instance of AfterQuery being named; it doesn't demonstrate the absence of other named partners elsewhere in the report. I cannot access the report to check this myself, and even the excerpt provided doesn't logically establish exclusivity.
To check: A full read of the NVIDIA Nemotron 3 Ultra technical report's acknowledgments/data-partner sections.
- 5¶
NVIDIA's report states the AfterQuery task distribution was selected for shared latent structure with GDPval along four named dimensions.
“"We then constructed a training distribution from AfterQuery (AQ) tasks that share important latent structure with GDPval, including file-grounded reasoning, professional deliverables, multi-step analysis, and judged final outputs.”
assertion · unclear
unverifiable · low confidence — Presented as a direct quote from NVIDIA's report; I cannot verify its authenticity or accuracy, and the underlying report postdates anything I can check. The idea of matching training-data 'latent structure' to a benchmark's task structure is a familiar, established practice in ML (train/test distribution alignment), even if this particular instance is unverifiable.
To check: The original NVIDIA Nemotron 3 Ultra technical report text, compared verbatim to this quotation.
- 6¶
The light SFT warmup on AQ rollouts was intended to transfer a strong model's workflow priors to the student before pivot RL.
“The goal of this step was to transfer the strong model's workflow priors for GDPval-like tasks to the student.”
assertion · unclear
unverifiable · low confidence — This is standard distillation/SFT-warmup rationale (teacher-to-student prior transfer is a well-known technique), so the mechanism described is plausible in general ML terms, but I cannot verify this specific quoted passage or its results without access to the source report.
To check: The NVIDIA technical report's methods section describing the SFT warmup stage.
- 7¶
GDPval covers 44 occupations and 1,320 tasks (220 public), authored by professionals averaging 14 years of experience.
“It spans 44 occupations across the nine largest sectors of US GDP, with 1,320 tasks (220 of them open-sourced) drawn from the real work of industry professionals who average 14 years of experience.”
quantity · unclear
consistent · medium confidence — This matches my understanding of OpenAI's GDPval benchmark (announced 2025): it covers dozens of occupations across major GDP sectors, includes on the order of ~1,300 tasks with a public subset, and tasks were built with experienced professionals. I recall the broad shape of these figures but cannot independently verify each exact number given the recency of the benchmark relative to my training.
To check: OpenAI's GDPval paper/announcement, which specifies occupation count, task count, open-source subset size, and professional experience statistics.
- 8¶
GDPval-AA v2 is claimed to carry the largest weight of any evaluation in Artificial Analysis's Intelligence Index.
“Artificial Analysis maintains a public leaderboard version, GDPval-AA v2, that scores models on the open tasks and is now the highest-weighted evaluation in their Intelligence Index.”
assertion · unclear
unverifiable · low confidence — I know Artificial Analysis as a real, active third-party benchmarking site with an 'Intelligence Index' composite score, but I have no knowledge of a 'GDPval-AA v2' leaderboard or its current weighting — this level of detail and versioning is likely newer than my training data covers.
To check: Artificial Analysis's published Intelligence Index methodology page listing per-benchmark weights.
- 9¶
GDPval-AA evaluation runs models as agents in a shell-and-web sandbox through the Stirrup harness.
“Models solve the tasks agentically, working in a sandbox with shell + web access via the Stirrup harness.”
assertion · unclear
unverifiable · low confidence — I have no knowledge of a harness named 'Stirrup'; sandboxed shell+web agent evaluation is a common established pattern (similar to other agentic benchmarks), but this specific tool name is outside anything I can confirm.
To check: Artificial Analysis's or the harness maintainer's documentation naming and describing 'Stirrup.'
- 10¶
Scoring is blind pairwise LLM judging by a rotating three-model panel, fit to an Elo scale where human expert work sits at 1,000.
“The resulting deliverables are compared in blind pairwise matchups, each graded by a judge sampled from a rotating panel of three frontier LLMs, and those results are fit to an Elo scale anchored to human expert work at 1,000 Elo.”
quantity · unclear
unverifiable · low confidence — Blind pairwise LLM-judged comparisons converted to Elo is an established evaluation pattern (used elsewhere, e.g. Chatbot Arena-style methods), so the mechanism is plausible, but I cannot confirm this specific leaderboard's panel size, rotation scheme, or anchor point.
To check: Artificial Analysis's published GDPval-AA v2 methodology documentation.
- 11¶
PivotRL filters training states down to turns whose sampled next actions yield mixed verifier outcomes, discarding saturated turns.
“It keeps only the turns where the sampled actions produce mixed outcomes—some pass, some fail—and discards turns that are already uniformly solved or uniformly failed.”
assertion · unclear
unverifiable · low confidence — This is a coherent, plausible RL design — filtering training states by outcome variance is a well-known idea in curriculum learning and active learning (e.g., prioritizing high-uncertainty examples) — but I have no record of a 'PivotRL' paper (Yi et al., 2026), which is dated after my training cutoff.
To check: The arXiv paper cited (2603.21383v1) describing PivotRL's turn-selection criterion.
- 12¶
PivotRL is reported to match end-to-end RL accuracy on SWE-Bench using roughly a quarter of the rollout turns.
“The intended benefit is lower rollout cost: on SWE-Bench, the paper reports accuracy comparable to end-to-end RL with about 4× fewer rollout turns.”
quantity · unclear
unverifiable · low confidence — A single-benchmark efficiency claim from a paper I have no record of and cannot check; the magnitude (4x) is plausible for a method that discards saturated/uninformative training turns, but plausibility is not confirmation.
To check: Table/figure in the PivotRL paper reporting SWE-Bench accuracy versus rollout-turn count for PivotRL vs. end-to-end RL.
- 13¶
Adding the SFT warmup lifts GDPval from 35.3 to 46.7, within 2.8 points of the specialized teacher at 49.5.
“On GDPval, warmup raises the MOPD result from 35.3 to 46.7, leaving Ultra only 2.8 points behind the office/workplace teacher.”
quantity · unclear
unverifiable · low confidence — These are specific ablation numbers attributed to an NVIDIA report I cannot access; the general pattern (SFT warmup substantially helps in-domain agentic tasks) is a well-established phenomenon in distillation literature, but I cannot confirm these particular figures.
To check: Tables 4–5 of the cited NVIDIA Nemotron 3 Ultra technical report.
- 14¶
BrowseComp shows a warmup gain of the same shape as GDPval, 33.0 to 44.4.
“BrowseComp shows the same pattern, rising from 33.0 to 44.4.”
quantity · unclear
unverifiable · low confidence — Same reasoning as the GDPval ablation figure: unverifiable specific numbers from an inaccessible report, though BrowseComp is a real benchmark (browsing/search agent tasks) I recognize.
To check: The same NVIDIA ablation table cited in the post.
- 15¶
On HLE without tools the warmup produces a negligible 0.4-point change, unlike the agentic benchmarks.
“HLE barely moves, from 26.3 to 26.7.”
quantity · unclear
unverifiable · low confidence — HLE (Humanity's Last Exam) is a real benchmark I'm aware of, and the described contrast (agentic/file-grounded benchmarks improve from workflow-prior transfer, a knowledge-heavy no-tools benchmark does not) is mechanistically sensible, but the specific numbers are unverifiable without the source report.
To check: The same NVIDIA ablation table cited in the post.
- 16¶
AfterQuery claims its own replication with a Nemotron 3 Nano student and pure on-policy distillation reached a +20.9% net win-loss margin over base.
“AfterQuery has similarly validated that on-policy distillation works well for improving models on GDPval-style tasks, reaching a +20.9% net win-loss margin over base with a Nemotron 3 Nano student with pure OPD.”
quantity · unclear
unverifiable · low confidence — This is the vendor's own internal, unpublished result presented with no methodology, baseline definition, sample size, or independent verification — a self-serving figure with no way for a reader to check it, and it references a 'Nemotron 3 Nano' model outside my knowledge.
To check: AfterQuery publishing the underlying experiment details (task set, judge setup, sample size) or a third party reproducing the +20.9% figure.
- 17¶
AfterQuery asserts its Office Agent dataset is structurally matched to GDPval tasks.
“AfterQuery's Office Agent tasks mirror GDPval task structure, with file-grounded inputs, multi-step analysis, and rubrics.”
assertion · unclear
unverifiable · low confidence — This is the vendor's characterization of its own product's fit to a benchmark it is trying to sell data for — plausible on its face, and reinforced by the (also vendor-adjacent) quoted NVIDIA rationale, but I have no independent way to assess the actual overlap between AfterQuery's dataset and GDPval's task distribution.
To check: A side-by-side comparison of sample AfterQuery Office Agent tasks against public GDPval tasks.
- 18¶
The post's reading of the recipe: PivotRL selects the training states and MOPD provides the learning signal at them — two complementary components rather than competing methods.
“In other words: PivotRL supplies the local "where should we train?" states, while MOPD (Multi-teacher On-Policy Distillation) supplies the teacher-student learning signal at those states.”
assertion · unclear
plausible · low confidence · novel — This is the author's own interpretive synthesis (flagged 'in other words') of how two techniques compose, rather than a claim directly sourced from NVIDIA's report; it's a coherent, tidy division of labor between state-selection and learning-signal that isn't obviously wrong given the pieces described, but I can't confirm it's how NVIDIA itself frames the interaction, and MOPD as a named technique is outside my prior knowledge.
To check: Whether NVIDIA's report itself uses this same framing, or whether ablations isolate PivotRL's state-selection role from MOPD's signal role.
The light SFT warmup on AfterQuery rollouts is what closes most of the student-teacher gap on GDPval, and it does so specifically for GDPval-like agentic work rather than for capability in general.stands on 5 unverifiable steps
premise · unverifiable — NVIDIA's report states the AfterQuery task distribution was selected for shared latent structure with GDPval along four named dimensions. · claim 5
premise · unverifiable — The light SFT warmup on AQ rollouts was intended to transfer a strong model's workflow priors to the student before pivot RL. · claim 6
evidence · unverifiable — Adding the SFT warmup lifts GDPval from 35.3 to 46.7, within 2.8 points of the specialized teacher at 49.5. · claim 13
evidence · unverifiable — BrowseComp shows a warmup gain of the same shape as GDPval, 33.0 to 44.4. · claim 14
evidence · unverifiable — On HLE without tools the warmup produces a negligible 0.4-point change, unlike the agentic benchmarks. · claim 15
- ¶
inference · ungraded — “For the office/workplace teacher, the AfterQuery tasks were chosen because they resemble GDPval: file-grounded reasoning, multi-step analysis, professional deliverables, and judged final outputs.”
- ¶
conclusion · ungraded — “On GDPval, warmup raises the MOPD result from 35.3 to 46.7, leaving Ultra only 2.8 points behind the office/workplace teacher.”
By training only at turns with mixed verifier outcomes, PivotRL reaches end-to-end RL accuracy at roughly a quarter of the rollout cost.stands on 1 unverifiable premise · partially graded
premise · unverifiable — PivotRL filters training states down to turns whose sampled next actions yield mixed verifier outcomes, discarding saturated turns. · claim 11
- ¶
inference · ungraded — “RL is then run locally at those retained "pivot" turns, using verifier rewards for functionally valid actions rather than exact matches to the demonstration.”
- ¶
inference · ungraded — “The intended benefit is lower rollout cost: on SWE-Bench, the paper reports accuracy comparable to end-to-end RL with about 4× fewer rollout turns.”
conclusion · unverifiable — PivotRL is reported to match end-to-end RL accuracy on SWE-Bench using roughly a quarter of the rollout turns. · claim 12