Case study · vord

vord: a judge the agent cannot talk round

A static analysis guardrail in Rust that decides whether an AI agent's write may reach disk, and a coding agent of its own held to the same judge.

Rust · tree-sitter · 24 languages · MCP · Claude Code hooksGitHub ↗

The problem

A coding agent proposes a change, decides the change is good and calls it done. Verification is usually another call to the same model: the agent marks its own homework. With AI writing more and more code, speed goes up and, without a safety net, so does instability: skipped layers, dependency cycles, user input that ends up in a shell.

Linters and analyzers join that conversation late. They answer “what is wrong with this code?” once the code exists, in the PR or in CI. I wanted to answer a different question, inside the agent’s loop and before the bytes reach disk: may this write happen?

Decisions

A deterministic judge, separate from the writer. vord hook plugs into the agent’s edit loop (Claude Code, Codex or a pre-commit) and checks every write against a policy. When it denies one, the reason goes back into the agent’s context so it can fix it. The judge is not a prompt: it is rules that read the syntax tree, so there is no talking it round.

Structure, not text. Every rule reads the neutral AST the tree-sitter parsers produce, not strings. A comment or a string that happens to look like a type test cannot produce a finding.

Architecture as an enforceable rule. Beyond security and complexity, vord applies zero-config hexagonal layering, domain purity from frameworks, Martin’s component metrics, SOLID and tactical DDD. The metrics are not new (they come from CodeQL, SonarQube or Martin); what is new is that they can be enforced, in one engine and across languages.

One binary, nothing else. No JVM, no server, no database and no network, unless you configure an LLM provider for the agent. It installs with a script, npm, Homebrew, cargo or Docker, and as a Claude Code plugin.

A person decides the exceptions. A blocked write can be escalated and approved once with vord hook approve, and every non-silent decision goes into a log you can query with vord hook audit.

What is built

  • More than 300 rules in 18 crates: OWASP and cross-file taint analysis, code smells, architecture, DDD, Rust, WordPress, mutation gaps and more.
  • 24 languages through tree-sitter, behind a shared AST.
  • vord agent: a coding agent that cannot approve its own work; it finishes when the analyzer sees nothing new, not when the model says so.
  • Flow coverage (vord flow add) to declare critical sequences that static analysis cannot rebuild on its own.
  • SARIF import from other tools (Oxlint, Ruff, Clippy) into the same quality gate, an MCP server and compliance reports.

What it does not do

vord has no type checker and no borrow checker. A write that passes every rule can still fail cargo build or tsc. It complements the compiler and the tests rather than replacing them: the agent should still run them before calling anything done.

What I learned

The useful question was not “how do I make the model more careful?” but “what infrastructure makes carelessness go nowhere?”. Putting the boundary outside the model is what makes autonomy auditable.

Next caseokf-mcp: agent memory you can audit →