Cantina open-sources Apex Flash-1: a 321B-param security model that nearly matches Claude Opus 5 High — for $2.38 instead of $74.68
Cantina Security, a startup that hunts software bugs for bounties, announced Sunday it is open-sourcing Apex Flash-1, a 321-billion-parameter model trained to read code, wield tools, and verify its own exploits — and its first published numbers suggest specialized security agents are about to get a lot cheaper to run.

Cantina Security spent the last few years finding other people's bugs. Sunday, it gave the tooling away. Co-founder and CEO Harikrishnan "Hari" Mulackal announced on X the release of Apex Flash-1, an open-weights model fine-tuned specifically for vulnerability research — the company's first public model release. The weights are live on Hugging Face for anyone to download and run.
The company behind it is not new to the game. Mulackal previously worked on the Solidity language and compiler at the Ethereum Foundation and co-founded the security firm Spearbit. Cantina emerged from stealth in July with an $8 million round led by Framework Ventures, bringing its reported total funding to $16.5 million, according to SiliconANGLE. The open model extends the strategy from selling security work to selling the whole stack: evaluations, data, harnesses, and now the weights.
Trained to verify its own exploits¶
Apex Flash-1 is a reinforcement-learning fine-tune of GLM-5.3-Flash, released under the MIT license and developed with partner Yeta. The model card lists 321 billion total parameters and says the checkpoint retains the base model's image-text-to-text architecture, though the published evaluation covers text-based security tasks only.
The training philosophy is what makes it interesting: rather than asking a general-purpose model to find bugs, Cantina trained a specialist to run the full loop — read the code, use tools, pursue an exploit, and verify the effect against a running target in production-like software and protocol environments. It is positioned as a worker model that a larger agent directs, and Cantina recommends running it through an agent harness such as Codex. Notably, the company does not yet offer hosted, token-priced access — the weights are yours to serve.

The numbers, with caveats¶
Cantina's model card reports an evaluation of 60 tasks built from 20 held-out vulnerability cases, each with guided whitebox, focused whitebox, and focused blackbox views. Targets ran in isolated environments, and verifiers checked the final target state. The headline scores: Apex Flash-1 solved 40 of 60 tasks (66.7% pass@1), the untuned GLM-5.3-Flash solved 36 (60%), and Claude Opus 5 High solved 43 (71.7%).
Then comes the part Cantina really wants you to see — the bill. Cantina estimates the 60-task run cost $2.38 for Apex, against $4.56 for the base model and $74.68 for Claude, using provider pricing. That's the pitch: near-frontier security capability for about 3% of the cost of renting a hosted frontier model.
The asterisks are worth reading twice. These are company-reported results from a small, company-designed test set — not an independent benchmark, and not evidence the model finds vulnerabilities at this rate in production. Each model ran the set once. Cantina itself acknowledges the limits, and has promised harder multi-step investigations and public benchmark results as they are validated. Mulackal also criticized public cyber evaluations as poor proxies for the work customers actually need — which cuts both ways, since it means outsiders can't easily compare his numbers against a neutral common test either.

An economics problem¶
Mulackal's framing is that security work is becoming an economics problem: organizations need to find and verify more vulnerabilities without spending as much on each investigation. In his post he claimed Cantina has earned $1 million in bug bounties and ranks first on HackerOne's US business leaderboard for 2026, and said its security harnesses process trillions of tokens a month. Those are company claims, not audited figures — but they sketch the worldview: a full-stack security company where the model is one component of an automated investigation pipeline.
The timing lands in a broader shift in how the industry talks about AI's offensive potential. In a September 10 report, Anthropic described AI-assisted "exploit foundries" — workflows for vulnerability research and exploit development — and Mulackal cited it in his post as evidence that security work is being automated. Anthropic's account documents its own investigations of observed misuse, not validation of Cantina's model.
One more detail for the careful: alongside the standard release, Cantina also published an experimental derivative called Apex Flash-1 Abliterated, described as having modified refusal behavior. The reported evaluation applies to the standard model, not that variant. Releasing an open-weights vulnerability-research model with an explicit no-refusal sibling is a choice that will get scrutinized — and Cantina has made it in public, with the weights, the training recipe, and the caveats all visible.
Why it matters¶
The dual-use question hangs over the whole release. A cheap, open security agent that verifies its own exploits is exactly what defenders want for continuous, affordable testing — and exactly the kind of capability offensive actors also benefit from. Cantina's answer is transparency: publish the weights, the eval, and the limitations, and let the economics of defense finally catch up with the economics of attack. Whether the 66.7% holds up outside Cantina's own environments is now a question anyone with a GPU can go and answer. That, more than the headline number, is the point of open weights.