I build local-first AI systems, evaluation tools, and practical engineering products.
The strongest thread across the work is simple: make AI agents cheaper to run, easier to inspect, harder to fool, and safer to trust with real code and real decisions.
- Memento Mori Jester: a local MCP/CLI sidecar that reviews agent plans, commands, diffs, and final claims before risky work ships.
- AgentLedger: a local-first black box recorder for AI coding agents, with command evidence, repo state, audit reports, and handoff bundles.
- TokenSquash: a measurable prompt/reply codec and evidence harness for testing whether repeated agent traffic can be shortened without losing meaning.
- RepoMori: machine-readable repository packs for AI agents and local tools, built around exact source recovery and compact context.
- ManifoldGuard: a reference-bounded output regulator for checking candidate outputs against supplied semantic and relational structure.
- The Gauntlet: a local-first paper and theory stress tester with transparent rule-based verdicts and source-grounded reports.
I am converging the public work into a local AI agent control stack:
- memory and context: RepoMori
- execution evidence: AgentLedger
- safety review: Memento Mori Jester
- token and usage pressure: TokenSquash and Tokometer
- output grounding: ManifoldGuard
- evaluation and stress tests: The Gauntlet, The Marked Bench, and the consequence-agent benchmark work
The private work continues this same direction through AIOS, MiddleOut, and consequence-memory systems.
- Tokometer: a local usage gauge for Codex token burn, rate limits, history, alerts, and exports.
- The Marked Bench: a versioned contradiction-detection benchmark for AI reasoning evaluation.
- Rolefit CV: a local-first CV and job-fit assistant that keeps claims grounded in real evidence.
- ChatP2P: peer-contributed AI compute with signed nodes/jobs, verified results, and credit-based coordination.
- Motion-TimeSpace: a research workspace connecting physics thinking with reproducible computational tools.
- Local-first tools that users can inspect and run themselves.
- Evidence over vague claims.
- Practical release gates, tests, and reproducible demos.
- AI systems that remember what happened, show their work, and admit limits.
- Collaboration on AI agent tooling, benchmarks, safety, observability, and local-first products.
- Technical review of the public tools above.
- Contract or freelance work where reliability, clarity, and shipped artifacts matter.


