Appier's SMITH trains AI agents to forge their own tools — a 4B model beat a 30B one at it
Appier announced today that its paper on self-forging AI tools has been accepted at NeurIPS. The SMITH framework trains agents to build and use tools in one reinforcement-learning loop — and the tools written by a 4B model beat those from a 30B baseline.

Appier announced today that its paper “Joint Optimization of Tool Creation and Use for Large Language Model Agents” has been accepted at NeurIPS. The paper introduces SMITH — Schema-grounded Multi-task Iterative Tool Honing — a reinforcement-learning framework that trains an AI agent to both create tools and use them inside a single training loop.
The problem it attacks is the one every agentic deployment eventually hits: models are only as useful as the tools they can call, and someone — usually an engineer — has to build those tools. Existing tool-creation methods hand the job to a separate model from the one that uses the result, so the builder never gets feedback on whether a tool is clearly described, runs reliably, or can actually be invoked. SMITH closes that loop: when a tool’s description is vague, its parameters misdesigned, or its code fails, the failure feeds straight back into training.
Build it, use it, keep what works#
SMITH trains easy to hard. The agent first learns a method from 4 simple examples and writes tools from them, then faces 16 harder problems it has never seen. Only tools that solve new problems survive, joining a shared library that multiple agents can draw on — and better tools replace weaker ones over time. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question), with three separate reward axes catching schema, code, and outcome failures independently, so each failure mode teaches its own lesson.

A 4B model out-built a 30B one#
Three results stand out. First, a Qwen3 model of about 4 billion parameters trained with SMITH wrote tools that outperformed every other method in the study on unseen tasks — including a baseline in which a roughly 30-billion-parameter model built tools on the fly. On 13 procedural-reasoning tasks with exact verifiers, the 4B model reached 79.8% macro-average accuracy, the best score in the evaluation. Second, the tools travel: written by the small model, they lifted a 350-million-parameter lightweight model and also boosted larger models — which makes it cheaper to divide work across agents of different sizes.

The 32× efficiency play#
Third is the efficiency number. Because SMITH compiles repeated reasoning into callable tools, average output fell from 3,206 tokens with conventional step-by-step reasoning to about 100 — roughly a 32-fold drop — while task performance held. For anyone running agents at scale, that is the difference between a rounding error and a budget line.
“Humans turn their problem-solving experience into tools, so they never have to start from scratch. AI agents are now evolving in the same way,” said Dr. Chih-Han Yu, Appier’s CEO and co-founder. “This research shows that agents can learn to build tools, continuously refine them, and share proven tools across models of all sizes, making multi-agent collaboration more efficient and scalable.”
Chieh-Yen Lin, a research scientist at Appier, described the mechanism: “When SMITH trains a model to use tools, it sees only the tool’s description and parameter specifications, not the underlying code.” That constraint turns description clarity into direct training feedback.
What to watch#
Some proportion is in order. This is a company-announced research result — Appier is a Tokyo-listed adtech and martech vendor (TSE: 4180) — and the numbers come from 13 procedural-reasoning tasks with exact verifiers, a controlled setting rather than the open web. A NeurIPS acceptance is peer review’s stamp of interest, not a product launch. Still, the direction matches where the whole agentic stack is heading: tools agents write for themselves, verified by execution, shared across fleets. Appier frames the payoff in enterprise terms — repetitive operations like converting financial metrics, processing data, querying reports, and routing support cases could stop being rebuilt by hand every time and become proven, callable tools instead.
Sources
- PR Newswire (Appier) — “Appier Research Accepted at NeurIPS: AI Agents Learn Not Only to Use Tools, but to Build Their Own” (September 30, 2026)
- Appier — press release: “Joint Optimization of Tool Creation and Use for Large Language Model Agents” accepted at NeurIPS (September 30, 2026)
- arXiv — “Joint Optimization of Tool Creation and Use for Large Language Model Agents” (arXiv:2608.24571, August 25, 2026)