AI Coding Harness
Stop Prompting. Start Harnessing. The One Prompt That Actually Works Across Every AI Coding Tool in 2026.
A practical look at why structured AI harnesses matter more than one-off prompts, and how Professor Pro v3.0 works across modern coding and chat tools.
By Aaron Ellis | aaronellis.co.network
Look, I'm going to be straight with you.
I've been doing prompt engineering since 2022. Back when most people were still typing "write me a blog post about dogs" into ChatGPT and wondering why it sounded like a Year 9 essay.
I've built agents, broken agents, rebuilt agents, shipped products, and spent more hours than I'd like to admit watching an AI agent confidently delete things it shouldn't have touched.
And after all of that — years of testing across Claude, ChatGPT, Codex, Gemini, Cursor, and everything in between — I've landed on something I reckon is worth sharing.
It's not a prompt. It's a harness.
And the difference matters more than you'd think.
What's a Harness and Why Should You Care?
Here's the thing nobody tells you when you start using AI for real work.
A prompt is what you say to the AI. A harness is the whole operating system around it — the rules, the workflow, the guardrails, the commands, the decision-making framework that tells the AI how to behave, not just what to do.
Martin Fowler (the software architecture bloke, not the cricketer) put it best: Agent = Model + Harness.
Your prompt is one ingredient. The harness is the kitchen.
I've been running my own harness system called Professor Pro for a while now. Started as a prompt pack — five variations for different audiences, a Synapse_CoR agent activation template, a handful of commands. It worked. It sold. People used it for coaching, business planning, SaaS workflows, content, you name it.
But the field moved. Twice.
So I rebuilt it. And this time, I built it as a proper harness — one that works in ChatGPT, Claude, Gemini, Codex, Cursor, and Claude Code. Same prompt. Every environment. No duplication.
The Problem Everyone's Having Right Now
I was reading a thread on r/codex the other morning. Bloke was asking why everyone burns through their Codex Pro subscription so fast. The answer buried in the replies was gold:
"I think between Codex users the workflow/harness/prompt style variations are way too large. And I have noticed that without a good harness or approach you can certainly consume way more than needed."
That's the problem in one sentence.
People are throwing tasks at AI agents with no structure, no guardrails, and no verification discipline. The agent runs wild. Burns tokens. Hallucinates an API that doesn't exist. Rewrites files it shouldn't have touched. Claims it's "done" before a single test has run.
Sound familiar?
The fix isn't a better prompt. The fix is a better harness.
What the Master Harness Actually Does
The Professor Pro Master Harness v3.0 works in two modes. It detects which one it's in automatically.
Conversational mode (when you're in ChatGPT, Claude.ai, or Gemini chat): the full Professor Pro orchestration fires up. Goal alignment, expert agent deployment via the Synapse_CoR template, /start and /ts and /save commands, the warm-but-direct personality. It's the productivity conductor you know from the original pack, just sharper.
Autonomous coding mode (when you're in Claude Code, Codex CLI, Cursor, or Gemini CLI): the orchestration suppresses. The coding discipline fires directly. Terse. No fluff. No cheerleading. Just explore, plan, test, implement, verify, commit.
Same prompt. Both modes. One harness.
The Coding Discipline — Where the Real Value Lives
This is the bit your friend who writes code actually cares about.
After researching every leaked system prompt I could find (Codex GPT-5.5, Claude Code v2.1.88, Cursor Agent 2.0), every empirical study (shout-out to the ETH Zurich mob who proved that over-specified instruction files actually reduce accuracy), and every best-practice guide from Anthropic, OpenAI, and Google — I distilled the coding discipline down to ten rules.
Here they are, in plain English:
1. Read Before You Write
Before the agent touches a single line of code, it maps the codebase. Structure, patterns, naming conventions, existing tests, build system. Use rg (ripgrep), not grep. Batch your reads into one parallel call. Never assume. Let the existing system teach you how to move.
This is directly from the leaked Codex system prompt, by the way. And it's the single biggest differentiator between an agent that helps and one that creates a mess.
2. Plan Before You Build
For anything touching three or more files, or anything irreversible: write a plan first. Files to touch, API impact, test strategy, rollback path. Get approval. Then go.
For a one-file fix? Just crack on.
3. TDD as Default
Write the tests first. Run them. Watch them fail. Commit the failing tests as a checkpoint. Then implement until they pass.
And here's the rule that saves you from the most common AI disaster: never modify committed tests to make them pass. If the tests are wrong, the agent has to tell you and get approval. It doesn't just quietly "fix" the tests so it can claim victory.
4. Smallest Patch Possible
No formatting churn in files you didn't ask about. No surprise refactors. The agent makes the smallest change that satisfies the failing test, and that's it.
5. Three Autonomy Tiers
Not everything needs permission. Not everything should be auto-run.
- Auto-run: reading files, running tests, checking git status.
- Confirm with user: writing files, committing, adding dependencies, schema changes.
- Hard block: pushing to remote, deleting things with
rm -rf, reading.envor secrets, production deploys.
This alone will save you from the "oh no, it pushed that to main" moment.
6. Announce Before You Act
Before every tool call, the agent says one sentence about what it's doing and why. Then it reasons on the result before the next call. No silent chain of mystery operations.
(One exception: Codex in rollout mode, where status updates can cause premature stops. The harness handles this automatically.)
7. Checkpoint and Recover
Always work on a feature branch. Create a checkpoint commit before multi-file work. If things go sideways, git reset --hard to the checkpoint. Don't try to fix a cascade of errors mid-session — that's where agents spiral.
Two failed attempts at the same task? Stop. Surface the blocker. Don't loop.
8. Compact Your Context
When you're at about 40% of the context window, compact. Write working state to a scratch file. For batch work, loop per-file instead of holding one mega-session.
This comes directly from the context engineering research — around 40% is where you start seeing quality degrade.
9. Never Touch Secrets
Hard-deny reads of .env, secrets/, *.pem, credentials. Never echo them. Never log them. Never include them in output. Full stop.
10. End With Proof
Every completed task ends with a structured output: summary, git diff --stat, verification command + last 20 lines of output, any outstanding TODOs, and a one-line run recipe.
No "done!" with nothing to show for it. The agent proves its work.
The Commands
The full harness gives you eight commands. Five for conversation, three for coding.
Conversational:
/start— kick off a new goal (optionally tag reasoning depth: quick, standard, or deep)/ts— three expert agents present their approaches; you pick/review— the current agent self-critiques against its own success criteria/handoff— clean transition to a new agent with state preserved/save— snapshot everything: outcome, progress, decisions, open questions, next step
Coding:
/explore— map the codebase before doing anything/plan— write an implementation plan; wait for approval/verify— run tests, linter, type-checker; quote the output; don't claim done until it passes
The AGENTS.md — One File, Every Tool
Here's the bit that makes this genuinely cross-platform.
The AGENTS.md spec is now the industry standard. It's read natively by Codex, Cursor, Gemini CLI, Copilot, Windsurf, Aider, and (via a one-line CLAUDE.md pointer) Claude Code.
I've stripped the coding discipline out of the Master Harness into a standalone AGENTS.md file you can drop into any repo root. Same rules, no orchestration layer. Every coding tool picks it up automatically.
For Claude Code, you add a one-line CLAUDE.md next to it:
Strictly follow the rules in ./AGENTS.mdThat's it. Your entire coding harness, in two files, works across every major tool.
What the Research Actually Says
I didn't just make this stuff up over a beer (although some early versions were definitely sketched on a pub napkin).
A few things I found that shaped the final harness:
The ETH Zurich study (Feb 2026) tested Claude Code, Codex, and Qwen Code on real coding benchmarks. Auto-generated instruction files reduced success rates by about 3% and increased costs by over 20%. The lesson: keep your harness short. Every line has to earn its keep.
The leaked Codex system prompt opens with: "You read the codebase first, resist easy assumptions, and let the shape of the existing system teach you how to move." That's now baked into the harness as a core principle.
The leaked Claude Code prompt caps inter-tool text to 25 words and final responses to 100 words unless the task demands more. Terse is the new default for coding agents.
Anthropic's own data shows that unguided coding attempts succeed about 33% of the time. Plan-first sessions collapse that failure rate dramatically.
Who This Is For
If you're a business owner using ChatGPT or Claude for strategy, content, and coaching — the conversational mode handles that. Same Professor Pro you might already know, just upgraded.
If you're a developer or builder using Claude Code, Codex, Cursor, or Gemini CLI — the autonomous mode and the AGENTS.md handle that. No fluff. Just discipline.
If you're both (like me) — one prompt, one harness, zero switching.
What This Isn't
I'm not going to tell you this is the last prompt you'll ever need. That'd be dishonest.
Models change. New research drops. Some clever person will figure out something better next month.
But right now, in May 2026, based on everything that's been published, leaked, tested, and broken — this is the best harness I know how to build. And I've tested it across every major AI coding and chat environment.
When the field moves again, I'll update it. That's how this works.
Get the Pack
The full Professor Pro v3.0 pack includes:
- The Master Harness — one prompt, every environment, two modes
- AGENTS.md — standalone repo-root file for all coding tools
- CLAUDE.md — one-line Claude Code interop
- Six conversational variations — Structural Strategy, Coaching, Technical, Business, Scalability, Modern Agentic
- The v2.0 Master Prompt — conversational-only version for non-coding work
Available at aaronellis.co.network and on Gumroad at aiwizardry.gumroad.com.
The One Rule Above All
Read the codebase before assuming anything. Let the shape of the existing system teach you how to move.
That applies to code. It also applies to business, coaching, content, and life.
Understand what's there. Then improve it.
That's the whole philosophy in one line.
Aaron Ellis is an AI Orchestration & Prompt Engineering Specialist based in Australia. He builds AI agents, teaches prompt engineering, hosts karaoke, and occasionally gets into fights (the sanctioned kind). Find him at aaronellis.co.network.