Maggy, palgu>'s open-source AI engineering harness.
From the team behind palgu>: Maggy is our open-source harness that turns Claude Code, Codex, Kimi, and Gemini CLI into a test-enforced, quality-gated workflow, with cost-aware routing across 13 model tiers, cross-session memory, and a plugin system. Run it locally, own every line. Point it at palgu> when your team needs a governed gateway.
- Run your own multi-model routing locally, your keys, cheapest capable model per task.
- Add a real engineering harness to your CLI, TDD, gates, memory, agent teams.
- Self-host and own it end to end, open source, no account, no gateway required.
Two parts, one repo
Start with Bootstrap in 30 seconds; add the Maggy server when you want routing, protocols, and the dashboard.
An installable config pack (67 skills, hooks, rules, templates) that drops into ~/.claude/ and applies on your next session. Roughly 30-second install; works with Claude Code, Codex, Kimi, and Gemini CLI.
An optional local FastAPI server plus web dashboard adding 13-tier routing, skill protocols, the Cortex code graph over MCP, Polyphony isolation, and plugins. maggy serve → localhost:8080
Core systems
Everything Maggy adds, in one place.
Every task goes to the cheapest model that can actually do it.
- Each message is scored 1-10 for complexity and risk (the "blast score") and classified by task type, all locally, by a free Qwen3 classifier, so triage costs nothing.
- The score selects a tier: T0 Qwen3 (local) for classification and bulk ops, climbing through DeepSeek, Kimi, Gemini, Grok, and Codex, up to Claude Sonnet/Opus (T11 to T12) for architecture and security review. Roughly 80% of real work lands on the cheap-but-strong middle tiers.
- Budget-capped with auto-demotion: when an org nears its cap, routing steps down to cheaper tiers instead of blocking. Fatigue-aware and cascading, a failed call falls back down the chain.
Deterministic, intent-matched workflows, not just prompts.
- Protocols are YAML files (maggy/skills/protocols/). When your intent matches one ("push to git", "ship a feature"), Maggy runs the steps in order with real gates.
- Example, git-push: lint, typecheck, tests, stage, commit (with an AI-written message), push. Any failing step halts the protocol; nothing ships half-done.
- Drop a .yaml to add your own. 67 bundled skills cover Python, TypeScript, React, React Native, Flutter, Supabase, Stripe, Playwright, security, ADRs, and cross-agent delegation.
Tests tell you it passes; Telos tells you it fulfils its intent.
- Telos scores work on an Intent Fidelity Scale: IFS = F1 x F2 x F3 across three planes, Conformance (does it meet the written spec), Validation (does it do the right thing), and Integrity (is it sound and safe).
- The score is multiplicative: a zero in any single plane collapses the total to zero. You cannot pass by acing two planes and ignoring the third.
- Ships as a plugin, so it runs as part of the pipeline rather than as a manual afterthought.
A queryable code graph any harness can use over MCP.
- Cortex builds a code-property graph with 10 edge types, cyclomatic-complexity metrics, FTS5 full-text search, and bidirectional traversal, all in a single SQLite database.
- It exposes 15 MCP tools, so Claude Code, Codex, or any MCP client can ask structured questions ("what calls this", "what would this change break") instead of grepping.
- Benchmarked against plain codebase-memory approaches: graph traversal beats flat RAG for "why" and "blast radius" questions.
Run multiple agents on one repo without file conflicts.
- Concurrent agent sessions each get a Docker-isolated workspace, auto-provisioned when a second session starts.
- Because each agent works in its own isolated checkout, parallel work never clobbers another agent's files, the classic failure mode of running multiple agents on a shared repo.
- Isolation modes let you choose how strict the separation is for a given run.
Memory that survives compaction and persists across weeks.
- Mnemos is task-scoped memory with a four-dimension fatigue model and typed checkpoints. It ingests Claude Code session transcripts, scores how "hazy" each session is, and auto-checkpoints before context is lost.
- When context compacts, Mnemos restores the typed checkpoint instead of making you re-explain the task, freeing tokens while keeping the thread.
- Engram is the long-horizon layer: it persists architectural knowledge across weeks and handles seven distinct "amnesia" types so decisions do not evaporate between sessions.
Stores why code exists, not just what it is.
- The intent-augmented Code Property Graph records the reasoning behind code, ReasonNodes and constraints, alongside the structure.
- It detects drift across six dimensions, flagging when an implementation has wandered from its stated intent.
- And it prevents duplicate implementations: before building something new, the graph can tell you it already exists.
A six-agent TDD pipeline with enforcement that does not depend on remembering.
- Six roles run each feature: Lead, Quality, Security, Review, Merger, and Feature, a real pipeline, not one model doing everything.
- Stop-hooks enforce TDD: tests must pass before a task is considered done. No green, no merge.
- Quality gates are enforced per file, max 20 lines per function, 3 parameters, 2 nesting levels, and non-trivial changes require an ADR (reverse-engineered from git history if one is missing).
Routing: 13 tiers, cheapest capable wins
Every message is scored 1-10 for complexity and classified by task type (locally, by Qwen3). Trivial asks stay free and local; hard architecture climbs to Claude. Budget-capped, with auto-demotion.
| T0 | Qwen3 (local) | Classification, triage, free bulk ops |
| T1 | Gemini Flash-Lite | Bulk extraction, pipelines |
| T2 | DeepSeek Flash | Docs, tests, scaffolding |
| T3 | Gemini Flash | Multimodal, vision, audio |
| T4 | DeepSeek Pro | Complex coding, refactors |
| T5 | Gemini CLI | Multi-file agentic coding |
| T6 | AGY | End-to-end (git + code + test) |
| T7 | Kimi | Long-context analysis |
| T8 | Gemini Pro Search | Deep research, 2M context |
| T9 | Grok | Competitor intel, reasoning |
| T10 | Codex | Bulk generation, security-sensitive |
| T11 | Claude Sonnet | Quality-critical code, debugging |
| T12 | Claude Opus | Architecture, security, ADRs |
Plugins: drop-in extensions
A simple plugin system: drop one in and it hooks into Maggy's events. Ships with several out of the box.
Turns shipped work into posts and publishes to LinkedIn, X, and Reddit, with a voice engine (plain-text, no em-dashes) and a comment-reply heartbeat.
GitHub, Asana, and Monday providers wire tasks and issues into the harness.
Telos ships as a plugin; write your own with a small plugin.yaml plus plugin.py.
What makes Maggy different
Most "AI engineering" tools are either an autonomous agent loop (Hermes-style) that replaces your CLI, or a thin wrapper that adds nothing. Maggy is a discipline and routing layer that augments the tools you already use.
| Capability | Maggy | Autonomous agent frameworks | Raw Claude Code |
|---|---|---|---|
| Multi-model routing | Yes: 13 tiers, cheapest capable model | Usually single model / BYO loop | No: one model for everything |
| Discipline enforced | TDD stop-hooks + quality gates + ADRs | Optional, prompt-dependent | None |
| Cross-session memory | Engram / Mnemos, survives compaction | Rare / vector-store bolt-on | Lost each session |
| Code intelligence | iCPG + Cortex MCP (why code exists) | Usually plain RAG over files | None |
| Parallel safety | Polyphony, Docker-isolated workspaces | Manual / conflicts on shared repo | N/A |
| Works with your CLI | Augments Claude Code / Codex / Kimi / Gemini | Replaces your tools with its own | n/a |
| Open source | MIT, 1100+ tests, self-hostable | Varies | Closed |
What it looks like
You: "review the auth middleware for timing attacks" → Blast score: 8/10 (security + architecture) → Routed to: Claude (Tier 12) → ADR gate: found docs/adr/0003-jwt-strategy.md → injected as context → Review runs with full architectural context You: "push to git" → Intent matched: git-push protocol → ✅ lint · ✅ typecheck · ✅ tests · ✅ commit (AI-written) · ✅ push
Maggy and palgu>
They fit together. Maggy is the harness, local, free for individual developers, picking the right model and enforcing engineering discipline. palgu> is the gateway, the governed, audited, multi-tenant layer a team routes all of that traffic through for budgets, policy, and compliance. Point Claude Code at either; use Maggy solo and free, add palgu> when a team needs governance. Neither locks you in.