Why this is blowing up now

The most interesting conversational-agent project of the past month doesn't use a language model. NPC-Forge (gioblu/NPC-Forge) is a deterministic, CPU-only framework for building conversational agents — with no machine learning and no LLM anywhere in the loop. You define your agent's brain as a JSON dataset of intents, templates, and tools, and the framework matches user input to those intents with its own parser, tracks sentiment, holds multi-turn conversations, and can call tools like run_in_terminal when an intent says so.

At the time of writing it sits at 252 GitHub stars — this is an idea story, not a numbers story. Its terminal assistant TERMy front-paged Hacker News as a Show HN titled “TERMy – A fast terminal assistant that does not use LLMs” (2026-09-04), and it turned up again this week in AI-trending roundups as the week's talk of Hacker News. The repository is alive: the latest commit landed two days before this article. It's AGPL-3.0, so you can read every line of the matching engine before you trust it.

Why is it worth your time? Three reasons. First, determinism: the same input always produces the same intent match, which means you can regression-test your agent — and the project ships a test suite to do exactly that. Second, cost: zero inference, zero API bills, nothing leaves your machine. Third, compatibility: every installed NPC is served through an OpenAI-compatible API, so your existing tools, scripts, and IDE extensions can talk to a deterministic agent exactly the way they talk to GPT.

Diagram: a user query flows into interlocking brass gears representing deterministic intent matching, fed by a stack of dataset cards, and a glowing response flows out into a terminal window
AI-generated illustration for AI Frontier Post — how deterministic intent matching works

What you'll need

  • Linux — this is an experimental release and Linux (or WSL2) is the only supported platform, per the README.
  • Python 3.8 or newer and git and curl.
  • A non-root user account. The installer refuses to run as root — I learned this the hard way in the sandbox and created a normal user. This is a deliberate safety choice, since installed NPCs can run shell commands.
  • About ten minutes. That's it: no GPU, no API keys, nothing to download — even TERMy's 24,070-intent dataset ships inside the repository itself.

Step 1 — Install NPC-Forge

Clone the repo and run the installer:

git clone https://github.com/gioblu/NPC-Forge.git
cd NPC-Forge
chmod +x setup.sh
./setup.sh

The script builds a virtualenv at ~/.local/share/npc-forge, pip-installs the framework and its Flask server stack, symlinks ~/.local/bin/npc-forge onto your PATH, installs a systemd user unit (npc-forge.service) that runs the gateway server, and seeds a tiny example NPC so you have something to poke at immediately. On my 2-core test VM the whole thing took a couple of minutes. Again: run this as your normal user, not root — the script exits immediately if it detects root privileges.

Step 2 — Meet the CLI: list NPCs and run the test suite

First, see what's registered:

npc-forge list

Out of the box you get the example NPC (80 templates, 500 intents, a 3.06 MB dataset). Next, run the built-in regression suite — this is the part that sets NPC-Forge apart from vibe-coded bots:

npc-forge test

37 tests ran in about 90 seconds: 36 passed, 1 failed, 2 marked as expected failures. The honest detail matters here. The failure was test_resolves_to_own_intent: the query “what does copy on write mean” matched a Python list-copy intent instead of the copy-on-write intent — a false-positive resolution. In other words, the matcher is a heuristic scorer, not magic, and the project's own suite catches it being wrong. Two tests are flagged as expected failures by the author. For a deterministic agent, having a test suite that honestly reports a wrong match is a feature, not a bug — it tells you exactly where your dataset needs more example utterances.

Step 3 — Install TERMy and talk to it

TERMy is the showcase NPC — a terminal assistant with a big authored dataset. Install it from the repo:

npc-forge install npcs/termy
npc-forge list

Now termy shows up: 100 templates, 24,070 intents, 158 vocabulary entries, 44.84 MB of dataset. That's the scale a hand-authored agent can reach — tens of thousands of intents, matched locally, in milliseconds. Talk to it:

termy "how are you"

I got a bordered response box: “I am fine yourusername, thank you! How is it going today!?” — the dataset uses <||username||> metadata that renders your shell username. One practical note: termy needs a real TTY. When I ran it from a non-interactive shell it hung waiting on stdin — in a normal terminal it just works. For scripts, pipe a file of prompts and add -y to auto-answer the permission prompts:

printf "how are you
" > test.termy
termy -y < test.termy

That printed the same response box — deterministic. Try a typo to see the matcher's second gear: I sent helo through the API and got back a probabilistic match with confidence 0.8 instead of an exact match, and a cheerful “hi there!”. The engine grades every match as exact, probabilistic, or rejected — and that grade is visible in every API response, which you'll see in the next step.

Step 4 — Serve it as an OpenAI-compatible API

This is the part that surprised me. NPC-Forge runs a gateway server on http://127.0.0.1:5000 that exposes every installed NPC as an OpenAI model. Start it:

npc-forge serve

That starts the npc-forge.service systemd user unit. (My sandbox has no user session bus, so I launched the underlying Flask app directly with python3 server.py from ~/.local/share/npc-forge — it's the same server the unit runs.) Now list your “models”:

curl http://127.0.0.1:5000/api/v1/models

You get an OpenAI-shaped list: example, smith, termy, each as {"id": ..., "object": "model", "owned_by": "user"}. There's also a native side: GET /termy/chat/ returns a compiled HTML chat interface for that NPC, and POST /api/chat/termy returns the raw engine result. Now the OpenAI endpoint:

curl -X POST http://127.0.0.1:5000/api/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"termy","messages":[{"role":"user","content":"how are you"}]}'

The reply is a proper chat.completion object — id, created, choices, finish_reason: "stop", even a usage block with prompt and completion token counts. Any OpenAI client library can point at this server and talk to a deterministic agent with zero code changes. Ask it something that triggers a tool:

curl -X POST http://127.0.0.1:5000/api/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"termy","messages":[{"role":"user","content":"show me the current date and time"}]}'

This time finish_reason is "tool_calls", and the message carries a function call: run_in_terminal with arguments {"command": "termy_say \"$(date)\"", "explanation": "...", "goal": "Tell the user the current date and time", "mode": "sync"}. A harness (the termy client, your own code) executes the command and feeds the result back. The native endpoint shows you the engine's full reasoning state — match status, confidence, sentiment counters (it tracks expletives, thanking words, interjections, and more), the matched intent's thinking text, tools, and the permission level (ask vs yolo) that decides whether a tool needs your confirmation.

Diagram: a small local computer with a lightning bolt serves API requests from a laptop, a desktop monitor, and a phone, replacing a remote cloud
AI-generated illustration for AI Frontier Post — the gateway replaces a remote LLM endpoint with a local deterministic one

Step 5 — Build your own NPC in one JSON file

This is the payoff: authoring an agent is writing a dataset, not training anything. Scaffold a new NPC:

npc-forge create gitbot

That creates ~/.local/share/npc-forge/npcs/gitbot/ with a config.json and a dataset/ directory. The dataset format (NDF 0.0) is plain JSON — an array of intent objects. I replaced the scaffold's dataset with a small git-status helper I wrote:

[
    {
        "category": "git_status",
        "input": [
            "what changed in my repo",
            "show me the git status",
            "any uncommitted changes",
            "git status please"
        ],
        "output": "Here is the state of your working tree:",
        "tools": [
            {
                "name": "run_in_terminal",
                "arguments": {
                    "command": "git -C /tmp status --short --branch",
                    "explanation": "Shows modified, staged and untracked files in the repo.",
                    "goal": "Report the git working-tree status",
                    "mode": "sync"
                }
            }
        ],
        "permission": "yolo"
    },
    {
        "category": "git_help",
        "input": ["how do i commit", "teach me to commit changes", "commit my changes"],
        "output": "Stage with git add, then record with git commit -m \"message\". Push with git push when you are ready.",
        "tools": [],
        "permission": "yolo"
    }
]

Each intent is a category, a list of example input utterances the matcher scores against, an output template, optional thinking text, a tools array, and a permission — yolo runs tools without asking, ask prompts for confirmation. Restart the gateway so it picks up the new NPC (npc-forge reboot), and query it through the same OpenAI endpoint:

curl -X POST http://127.0.0.1:5000/api/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"gitbot","messages":[{"role":"user","content":"what changed in my repo"}]}'

Back came finish_reason: "tool_calls", content “Here is the state of your working tree:”, and the function call run_in_terminal with my exact command, explanation, goal, and mode. That is a complete round trip: authored intent, deterministic match, structured tool call, OpenAI-compatible wire format. From npc-forge create to a working tool-calling agent: about thirty lines of JSON and one reboot.

On speed: once the server has warmed up (the first request loads the datasets and takes a few seconds), my requests against the 24,070-intent TERMy dataset answered in 0.01–0.05 seconds on a 2-core VM. That's the deterministic payoff — no queue, no token-by-token generation, no variance.

How it thinks — and where it breaks

Under the hood, the FlintParser engine tokenizes your input, scores it against every intent's example utterances (with typo tolerance), and returns the best match graded as exact, probabilistic, or rejected. Alongside that it runs a sentiment analyzer — every response I inspected carried counters for expletives, encouraging words, thanking words, interjections, and discouraging words — plus slot extraction, multi-turn conversation context, and personality phrase pools for completions, unknowns, and related-topic nudges. The project's docs also describe a hybrid escape hatch: when the parser returns rejected, the framework can escalate to an LLM for open-ended chit-chat, keeping deterministic speed for the queries your dataset covers and probabilistic flexibility for the rest. (That escalation path is documented, not something I wired up for this tutorial.)

The honest caveat is the failing test from Step 2. Deterministic does not mean infallible: “what does copy on write mean” matched a Python list-copy intent, confidently and wrongly. When the matcher is wrong, it's wrong the same way every time — which is exactly why the test suite exists. Authoring a good dataset is the real work here: every intent needs enough varied example utterances that near-misses resolve to the right place. TERMy's 24,070 intents didn't write themselves. Treat dataset curation the way you'd treat writing tests: it's the product.

When to use NPC-Forge — and when not to

Reach for it when determinism, cost, or privacy is the point: game NPCs with authored dialogue trees, CLI assistants, kiosk and helpdesk bots, offline or embedded agents, and any automation where an auditor needs to know exactly why the agent said what it said. A quick license note: it's AGPL-3.0, which has real copyleft obligations — read the license before you embed it in a commercial product.

Don't reach for it when the input space is genuinely open-ended — that's the LLM's home turf, and the framework's own docs concede it with the escalation design. Compared with Rasa, NPC-Forge is far lighter (no training pipeline, no ML ops) at the cost of less linguistic sophistication. Compared with hard-coded scripts, you get typo tolerance, sentiment tracking, multi-turn context, tool calls, and an OpenAI-compatible API for the price of a JSON file. And compared with any LLM agent: it's free, instant, offline — and it can't surprise you, which is either its superpower or its ceiling depending on the job.

One more piece of honesty: the author calls this an experimental release and says it's not yet production-ready. Linux-only for now, a small community, and the framework's maturity is closer to a very promising prototype than to infrastructure. Judge it as such.

The takeaway

In about ten minutes, with no GPU and no API keys, you installed a deterministic agent framework, regression-tested its matching engine, chatted with a 24,000-intent terminal assistant, served it to the world as an OpenAI-compatible API with real tool calls, and built your own tool-calling NPC from thirty lines of JSON. Every command above ran exactly as shown on a plain 2-core VM.

The deeper point is the one that got NPC-Forge onto Hacker News's front page: not every conversational agent needs a language model. When the domain is bounded and the behavior must be auditable, a scored intent match over a hand-authored dataset is the simpler, cheaper, faster tool — and now you know exactly how to build one.