
Vibe Coding is Dead.
Start Agent Engineering.
$/agt "ship the feature"$/agt "review the auth changes"$/agt "debug the flaky test"/agt picks the lightest setup that works, agents do the work, and you get one short card at the end.
One command in, one card out.
Type one command. It plans, builds and reviews, then ends on one short card. The plan, the findings and every agent report go to a file you open only if you want them.
$ /agt "refactor the dashboard sidebar to be collapsible"
recon: panes · 2 slices can build at once [agentille v3.0.0]
✓ collapsible sidebar on /dashboard · 6 files · PR #42 · 7m
verify: npm test → 48 passed
review: PASS — code-reviewer design-reviewer
Open: ~/.agentille/state/run-k7f2/report.md
Next: merge PR #42
When is it worth typing /agt?
Plain Claude is one developer. /agt is that developer plus a planner and a reviewer, brought in only when the task is big or risky enough to pay for them.
- A multi-step feature
/agt "add a dark mode toggle to settings, persisted per user, with tests" - Two independent parts
/agt "add a CSV export endpoint and a download button on the reports page" - A risky change
/agt "add webhook handling for subscription cancellations" - A review before merge
/agt "review the changes on this branch" - A bug across several files
/agt "debug why checkout totals are off by one cent" - Unsure of the scope
/agt --plan "migrate the auth pages to the app router"
What you get over a plain prompt
- A separate reviewer checks the work, not the model that wrote it.
- Each step runs on the model that fits it.
- Independent parts build at the same time.
- Auth, payments and data changes get a red-team pass.
Skip it for
- one-file edits
- typos and renames
- questions and explanations
- quick config changes
A plain prompt is cheaper there. Any run that is not solo costs more tokens than a plain prompt: what you buy is fewer wrong or unreviewed changes, not a lower bill.
Eight specialists. Always the right one.
agentille routes tasks to the agent built for them — not a generic chatbot doing everything at once. On UI work, design is a contract: one agent frames it, another builds it, a third reviews it.
planner
Goal-backward plan with parallel-step markers
plan-reviewer
Critiques the plan before any executor runs — coverage, parallel safety, scope. Opus for large or cross-cutting plans.
ui-prototyper
Designs the UI before the build — tokens, component anatomy, every state, anti-generic guardrails — as a blueprint the executor builds against. Uses your design skills when installed, its own taste when not.
executor
Headless implementation, atomic commits, adaptive PR/push. On UI work it builds against the prototyper's blueprint.
code-reviewer
Bugs, security, quality — read-only, severity-classified. Sonnet on small diffs, Opus when large or cross-cutting.
design-reviewer
6-pillar review, axe-core a11y per viewport, contrast + color-blind checks, AI-design-tell scan
adversary
Red-team tester. Writes tests to break the build, never source. Runs in the gauntlet formation.
security-reviewer
Secret leaks, injection, auth bypass — severity-classified
Tokens go where they earn the most.
Each model does exactly the job it's priced and capable for. No Opus on boilerplate. No Haiku on architecture. Fable only when failures prove it's needed.
Opus
Thinks hard so executors don't have to
- Goal-backward planning
- Architecture decisions
- UI prototyping (design blueprint)
- Design review
- Security review
- Large / cross-cutting code review
Sonnet
Fast, capable, ships the actual code
- All execution & coding
- Small-diff code & plan review
- Adaptive integration (PR / push / local)
Haiku
Dirt-cheap, runs at the edges
- Task classification, when the fast rules do not match
- Never writes or reviews code
Fable
Rare by construction, only on evidence
- Only after two failures at Opus max effort
- One per run, off above your weekly usage ceiling
- /agt --fable forces it on judgment roles
Hard work gets a shape.
A formation changes how the workers relate, not where they run. Each one spends tokens to buy something specific, and says what it costs.
duel
Two executors, two strategies. Tests, then a judge, pick one.
- buys you
- The better of two approaches when the right one is unclear
- costs
- ~1.8–2× build tokens
- runs
- Only when you ask
The loser is never pushed.
gauntlet
Build, attack, fix. An adversary tries to break the build.
- buys you
- Defects caught by tests before review
- costs
- ~1.3–1.6× build tokens
- runs
- Auto on auth, money or data changes with a test runner
The adversary writes tests only, never source. At most two rounds. The tests stay as regressions.
relay
A contract first, then coupled slices build in parallel.
- buys you
- Faster wall-clock on slices that share an interface
- costs
- ~+10–20% tokens over sequential
- runs
- Only when you ask
The contract is frozen while slices build.
/agt --formation duel|gauntlet|relay- Force one. If it doesn't fit the work, you get one honest line and a normal run.
See who's working, on what, for how much.
While /agt runs, a live band sits above your prompt: one row per agent with its model and effort, state, elapsed time and tokens. When the router escalates a role, the row says why. The band is drawn by code, so the model writes almost nothing mid-run.
/agt-ledgerand in the run's report file./agt-ledger- Print tokens per role for the run, finished agents included.
Only what needs you.
A run produces a lot of output. You see two things: flags while it runs, and one short card when it ends. Everything else goes to a report file.
report.md. Blocking findings (P0, P1) are fixed or flagged before the card can say ✓.Flags
Mid-run, no model call
- Read straight off agent results
- A revised plan, a FAIL or CONCERNS review
- A failed or skipped check
- An adversary that broke cases
- A pane waiting on you
- Each one also toasts
The card
End of run, eight lines at most
- What changed, in one line
- The verify command and its result
- The review verdict
- Only what needs you
- The path to the full report
- One next action
/agt-focus on|off- Agent flags above the band. On by default.
/agt-highlight on|off- Lit paths, versions and numbers in replies. On by default, remembered.
Parallel slices get their own pane.
Inside Herdr or tmux, when two or more slices can build at once, each slice's executor runs as a real Claude session in its own pane, named agt-<run>-<role>. You can watch every one, and talk to it. When its output is in, the pane closes. The mod opens and closes them: never focused, and only agt- panes.
gets a pane
The executor of each parallel slice, and the adversary.
stays a subagent
Everything else: the planner, every reviewer, the ui-prototyper, and any build with a single slice.
A pane is routed like a subagent. It starts as claude --agent agentille:agentille-<role> --model <m> --effort <level>, so it gets the same agent definition, model and effort. The first worker opens to the right of your session; later ones stack below it.
see it
Every worker is a full session in a pane you can read.
close it
Workers close once their output is in. The
agt-prefix is the line: panes you opened yourself are never touched.it says hi
Each worker pane has a small mascot: it waves hello, walks while it works and waves bye when done. Drawn by code, zero model tokens.
/agt-spawn "task"- Opens one extra Claude session in a new pane beside yours, on Sonnet by default. Typed only. It's yours, so the reaper leaves it alone.
What panes cost
On a two-slice test task (two runs per arm, alternating order), two Claude workers in panes compared with the same two workers as subagents:
panes ÷ subagents
- Fresh tokens, workers only
- 1.12×
- Fresh tokens, counting the lead
- 1.05×
- Counting cache reads, which bill at a fraction
- 1.55×
The gap is start-up context:
A pane opens at about 55k tokens against about 32k for a subagent: 1.12× the fresh tokens on this task. Not a saving. It buys a worker you can watch and talk to, which is why a single-slice build stays a subagent and panes open only for parallel work.
Small sample.
The old way vs. the one command
Spoiler: agentille does the chaining for you now.
Do it by hand
- 01
figure out the plan yourself - 02
pick the right model for each step - 03
remember to actually switch models - 04
branch + isolate the work - 05
build it - 06
remember to review it - 07
catch the UI regressions yourself - 08
open the PR by hand - 09
pray nothing conflicts
agentille
/agt "build the feature"
✓ done
(the orchestrator handled the rest)
Install the plugin.
Inside Claude Code. Takes two minutes. Then /agt is ready everywhere.
Requires Claude Code 2.1.287+ for the live band, routing and the pane mascot (the skills still work on older versions). Parallel panes need Herdr or tmux.
New to all this? Read the setup guide →
What you get
Orchestration benefits — not config files.
One-command orchestration
Type /agt "task" and it classifies, plans, dispatches agents and reviews. It ends on one short card, with the full detail in a report file.
Right model per job
Opus for planning and review, Sonnet for execution, Haiku for cheap classification. Tokens go where they earn the most.
Parallel-safe worktrees
Each executor branches off your current branch in its own git worktree. Atomic commits, then integrates adaptively — PR, push, or local branch.
Zero setup, your voice optional
Install and type /agt: no profile needed. Run /agentille-init once (five skippable questions) and every planner, executor and reviewer speaks your way.
Review built in
Code review on every change. Design review whenever UI is touched. Security reviewer available on demand. No extra step.
Works with your stack
No dependencies. Runs standalone, and automatically uses your design skills — impeccable, ui-ux-pro-max — when they're installed. Better with them, complete without them.
Audit it, fork it, ship it.
No black box. The entire orchestration is readable code. See exactly how agents are dispatched, models are routed, and your voice is applied.
github.com/hasuwini77/agentille
