Ivo open-sources Sage: the first legal AI model you can download, run, and fine-tune yourself
Contract-intelligence startup Ivo used its first Inscribe conference, on October 1, to do what no legal AI company has done: publish a free, open-source model post-trained for long-horizon contract work. It also previewed a benchmark that measures the thing current benchmarks ignore — an AI reviewer's judgment.

Legal AI is one of the most tightly guarded corners of the industry — everyone sells the service, nobody shares the model. Ivo just broke that pattern. At Ivo Inscribe, the company's first user conference for in-house legal and contracting teams in San Francisco, Ivo announced Ivo Sage: a free open-source model that attorneys, researchers, and developers can download, run, and build on top of — the first such release from a legal AI company.
Ivo is an AI-native contract intelligence platform built for enterprise legal teams — founded in New Zealand and headquartered in San Francisco — that turns contracts into what it calls actionable business intelligence. The company says Sage is built so legal practitioners can fine-tune it on their own datasets, adjust its rules and safeguards for specific use cases, and iterate on new approaches quickly, with every improvement open to the broader community.
How Sage was built#
Sage was built in partnership with River AI, which supplied technical and training-infrastructure support. Ivo post-trained DeepSeek V4 Flash on long-horizon contract work, using a mix of public data and synthetic data generated by real attorneys — then ran reinforcement learning on top.
The numbers, per Ivo: after RL, Sage went from meeting 70% of the pass criteria on the Legal Agent Benchmark (LAB) Contracts to 91%. The company says the model reached comparable quality to much larger frontier models, with higher token efficiency, at a fraction of the cost. Those figures are Ivo's own and have not been independently verified — the usual caveat for launch-day benchmarks.

A benchmark for judgment, not just edits#
The more interesting announcement may be the benchmark. Ivo previewed the Ivo-micro1 Contract Bench, built with data lab and research partner micro1, on the premise that existing contract benchmarks only grade the edits a model makes — while most of a reviewer's judgment happens before any edit: when to propose a change, when to leave acceptable language alone, and when to escalate to an attorney's judgment call.
The bench measures five dimensions of what a good attorney does: which issues to prioritize, when to show restraint, how to adapt to the specific deal, when to escalate, and how it follows the playbook. A public set of benchmark results is coming in the next few weeks.
Ivo shared four early findings from its research, and they are refreshingly unflattering to AI systems:
- Models are pushovers. When the right response to a counterparty's change is to accept it or leave it, models get it right 82% of the time. When the right move is to counter or reject, that drops to 46% — and they add a needed new point of their own only 23% of the time.
- Models don't escalate. On average they meet only 23% of the escalation criteria Ivo's expert attorneys set, despite explicit instructions and a dedicated escalation tool. GPT-6 Astra is the outlier at 71% — achieved by escalating nearly everything, which tanks its edit scores and puts it last overall.
- Models edit in broad strokes. Attorneys make 64% of their changes inline, averaging 92 characters. Models make 14% to 37% of theirs inline, rewriting whole blocks at 207 to 369 characters.
- Deal facts break playbooks. Models meet 51% of the criteria for applying a playbook's standard positions on average; when the right move depends on the deal's facts instead, that falls to 38%. Every model drops 9 to 15 points.

A full platform, too#
Ivo also announced the general availability of Ivo Collaborate, a full end-to-end contract intelligence platform for Fortune 500 legal teams. Collaborate reads incoming contracts and runs the whole workflow from intake through signature — moving routine agreements according to rules set by legal and flagging issues that need a person's judgment. A companion search layer, Intelligence, mines signed agreements so teams can see what the business has promised and is owed.
“The next leap for AI in contract and legal work won't come from bigger models. It will come from giving models the context of the work: the documents, the playbooks and the way a legal team operates,” said Min-Kyu Jung, Ivo's co-founder and CEO. “We're opening up our model so others can test it, adapt it and build on it. That's how we move closer to contract intelligence the whole industry can rely on.”
Why it matters#
An open-source model from a legal AI incumbent is a genuine first — and it lands at a moment when legal teams are being asked to trust AI with decisions that carry real liability. If Sage holds up outside Ivo's own evaluations, it gives every law firm and in-house team a starting point they can audit, adapt, and run behind their own firewall instead of renting a black box.
The benchmark is the longer play. A shared, public test of AI judgment in contract review — when to push back, when to hold back, when to ask a human — is something the field has been missing, and Ivo putting its early results out in the open is the right kind of transparency. The test now is whether other labs will run it — and whether Ivo publishes the full results set as promised in the coming weeks.