|

Harvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work

Harvey has launched Harvey Tenet, its first post-trained mannequin, as a research preview as of in the present day. Tenet is a Kimi K3 base post-trained with Fireworks by asynchronous reinforcement studying on long-horizon authorized work. The coaching corpus mixed artificial information, publicly obtainable authorized information, and human professional information. Harvey states no buyer information was used. Against the bottom K3 mannequin, Tenet completes nearly twice as many held-out duties on Harvey’s Legal Agent Benchmark (LAB) and 20% extra on LAB: Contracts, elevating all-pass fee by 9 and a couple of share factors respectively. Harvey studies state-of-the-art on LAB: Contracts and second place on LAB. The beneficial properties additionally transferred, untrained, to Mercor’s APEX Agents and Crosby’s Redline Bench. The said aim is twofold: construct frontier authorized intelligence on open-weight fashions, and provides legislation companies a path to personal their very own specialised fashions.

Is it deployable?

Not but, Harvey Tenet is a analysis preview introduced on August 20, 2026. Harvey has not printed weights, a mannequin card, or an API endpoint. The base mannequin is open-weight; Tenet itself is Harvey’s personal checkpoint, and the corporate says the work will transfer “from analysis to manufacturing” inside Harvey’s merchandise over time. What ships in the present day is the recipe, not the artifact.

  • Company tier: Enterprise solely. Access runs by Harvey’s platform, which is bought to law firms, mid-sized firms, and in-house legal teams. A lab with an RL stack might reproduce the strategy; coaching used roughly 150 NVIDIA B300 GPUs over two months.
  • Industries: Legal providers, company in-house authorized, personal fairness and funding banking (M&A diligence), plus regulated sectors the place contract quantity drives price — insurance coverage, monetary providers, healthcare, vitality.
  • Applications: M&A due diligence memos over datarooms, contract drafting, overview and redlining, structured extraction throughout as much as 10,000 paperwork, and precedent search over a agency’s collected data.

What the numbers say

Against the bottom K3 mannequin, Tenet completes nearly twice as many held-out duties on Harvey’s Legal Agent Benchmark (LAB) and 20% extra on LAB: Contracts, lifting all-pass fee by 9 and a couple of share factors respectively. Harvey studies state-of-the-art on LAB: Contracts and second place on LAB, utilizing base-model scores from Vals.

The extra fascinating result’s switch. Tenet additionally improves considerably on Mercor’s APEX Agents (company legislation) and Crosby’s Redline Bench — neither seen throughout coaching — whereas holding efficiency on data benchmarks together with LegalBench, CUAD, MAUD, and Scale’s PRBench. Agentic coaching didn’t erode textbook authorized reasoning.

Cost is co-optimized reasonably than traded away. Open weights lower cost per token; reward shaping that prefers shorter trajectories at equal high quality lowers tokens consumed. Harvey studies vital high quality beneficial properties at secure price.

How it was educated

Training used asynchronous reinforcement studying in sandboxed authorized environments constructed like LAB duties: a partner-style instruction averaging about 50 phrases, a shopper matter of key and peripheral paperwork, and an professional rubric of atomic go/fail standards — roughly 50 per activity, a whole lot on the excessive. A single rollout can exceed 1,000 turns.

Rollouts are graded by LLM-as-a-judge; ablations settled on Kimi 2.6. Reward combines the fraction of rubric standards glad, a holistic rely of authorized points solved, and an all-pass bonus. The coverage is optimized with GSPO utilizing a rank-64 LoRA over the complete K3 community, eight activity teams of eight rollouts per optimizer step, throughout ~1,750 environments and >10,000 rollouts per epoch. Fireworks co-built coach and rollout deployments on the kernel stage, with token-in-token-out and router replay, to maintain a big MoE numerically aligned throughout coaching and inference.


Three capabilities educated individually

Harvey group additionally post-trained specialist fashions that Tenet can path to as instruments or sub-agents:

  • M&A diligence: On LAB: Diligence, a single activity can traverse as much as 80M tokens; no baseline handed greater than 43.8% of standards. With Baseten, Harvey moved to a Recursive Language Model harness the place a root agent holds the dataroom in a REPL and delegates to sub-agents. A GLM-5.2 orchestrator alone reached 46.1%; post-training it in that harness through self-distillation reached 60.1%.
  • Review Table: With Applied Compute, a post-trained GLM-5.2 improved reply high quality by 3.6 factors and quotation high quality by 12.1 factors at roughly one-tenth the associated fee per cell, studying to abstain when a query doesn’t apply.
  • Firm data: With Engram, a Qwen3.8-27B mannequin research ~100M tokens of shopper issues into 1M tokens of structured data plus parametric reminiscence. Criteria go fee rose greater than 15%, tokens in accomplished trajectories fell 58%, and value per question dropped roughly 90% — 190.8 intelligence-per-token versus 129.3 for the perfect frontier configuration.

Marktechpost Independent Test Facts

Reality Check

Harvey Tenet — benchmark claims, verified

Audited: Harvey research preview, Aug 20, 2026 · and the Harvey X thread · Mode: default

78Inflation rating
19Claims
1Verified
10Self-reported
6Flagged
2Unverifiable

Nothing Harvey printed was contradicted. The rating is excessive as a result of Tenet seems on no public leaderboard — not Vals, not Artificial Analysis, not Mercor. Score formulation: (8 × 6 flags) + (15 × 0 contradicted) + (3 × 10 self-reported) = 78.

Claim desk

Claim Number Independent test Verdict
Completes ~2× extra LAB held-out duties than Kimi K3 base ≈ 2× Not on any public board Self-reported
LAB all-pass fee elevate +9 pts Same outcome said as “+82%” on X Flag F1
LAB: Contracts all-pass elevate +2 pts No public leaderboard exists Self-reported
State-of-the-art on LAB: Contracts SOTA Benchmark owned, run and graded by Harvey Flag F2
Places second on LAB #2 Vals #1 is Muse Spark 1.1 at 20.00%; Harvey-run, instrument delta by no means quantified Flag F3
Kimi K3 base on APEX Agents, company legislation 58.8% 58.8% (Kimi K3 Max) — matches precisely Verified
Tenet considerably beats K3 base on APEX Agents Harvey harness alone: 58.8% → 67.5% Flag F4
Beats K3 base on Crosby Redline Bench Not given Absent from the public leaderboard Self-reported
APEX v1 Big Law Associate held — a data benchmark, not the agentic board Not given Blind run commissioned from Mercor; not posted to the general public v1 board Self-reported
Holds on LegalBench, CUAD, MAUD “robust” Non-canonical metrics, utilized to all fashions Self-reported
PRBench arduous subset 36.0 → 36.8% Harvey calls it not statistically vital Self-reported
LAB: Diligence standards go fee 43.8 → 60.1% No public leaderboard Self-reported
Review Table price per cell ≈ 1/10 Baseline mannequin by no means named Flag F6
Firm Knowledge intelligence-per-token 190.8 Engram write-up; metric is Harvey’s personal Self-reported
“Our first post-trained open-weight mannequin” Business Insider: proprietary, in-house Flag F5
“Less than a fourth the price of main basis fashions” < 25% X solely; comparators unnamed Flag F7
≈150 NVIDIA B300 GPUs, 2 months, GSPO + rank-64 LoRA Unverifiable by building Not checkable
No buyer information utilized in post-training Unverifiable by building Not checkable

Flags defined

F1 · Denominator sportThe weblog studies +9 and +2 percentage points. The X thread studies the identical outcome as +82% and +22%. Both true; the social quantity sounds 9 occasions bigger.

F2 · Self-report as reality“SOTA on LAB: Contracts” is a win on Harvey’s personal benchmark. Harvey states there is no such thing as a public leaderboard for it and that every one scores are inner Harvey runs. LAB launched deliberately without a leaderboard.

F3 · Settings mismatchHarvey disclosed this plainly: Tenet ran in the usual public harness plus a end instrument carried over from coaching, whereas rival scores got here from Vals. The flag is about comparability, not concealment — Harvey by no means printed LAB with and with out the instrument, so its worth is unquantified. Harvey’s personal APEX figures present a harness change transferring naked K3 by 8.7 factors, and the LAB declare is a rank the place Vals’ leaders sit between 12% and 20%.

F4 · Settings mismatchOn APEX Agents, Tenet ran in Harvey’s inner bash harness whereas rivals used Mercor’s published numbers. Harvey discloses the harness lifts naked K3 from 58.8% to 67.5% — inside 0.1 pt of chief Fable 5 at 67.4%, earlier than any coaching.

F5 · Framing“Open-weight” describes the Kimi K3 base, not Tenet. No weights, mannequin card or API had been printed, but a number of retailers ran headlines calling Tenet itself an open-weight launch.

F6 · Denominator sport“Roughly one-tenth the associated fee per cell” is measured towards unnamed “strongest baselines,” with no serving config, precision or {hardware} given for both facet.

F7 · Denominator sport“Less than a fourth the price of main basis fashions” seems solely on X. The comparators are unnamed and checklist worth shouldn’t be separated from measured token consumption.

Credit the place dueHarvey had Mercor run APEX v1 blind, with out disclosing runs, duties or task-level scores again to Harvey — the strongest verification methodology within the publish, although it evidences data retention reasonably than agentic ability. Harvey additionally volunteered a null outcome on PRBench, documented its divergences from Artificial Analysis and Vals, and disclosed the harness impact in F4 that undercuts its personal APEX framing.

Reality Check by Marktechpost · verified 2026-08-23Default mode: vendor-only numbers accepted with a self-reported label. Scores change; re-verify earlier than citing.

Key Takeaways

  • Tenet is a post-trained Kimi K3 checkpoint, not a public open-weight launch — no weights, no API.
  • Gains transferred untrained to APEX Agents and Redline Bench, suggesting discovered conduct, not benchmark becoming.
  • Reward shaping on trajectory size made high quality and value enhance collectively as a substitute of buying and selling off.
  • The specialist stack — RLM diligence, Review Table, agency reminiscence — is the place the most important deltas landed.


Check out the TECHNICAL DETAILS here. Also, be at liberty to comply with us on Twitter and don’t neglect to hitch our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so forth.? Connect with us

The publish Harvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work appeared first on MarkTechPost.

Similar Posts