What is an Agent Harness? OpenClaw, Microsoft Autopilot and Meta Muse explained

- by

I kept hearing the word “harness” in AI circles and nodding along as if I knew what it meant. I didn’t. It usually turned up in the same breath as OpenClaw, which I also didn’t know, and then Microsoft announced something called Autopilot inside the new Copilot, and Meta launched an app called Muse, and it became clear that these were all somehow the same story. So I sat down with Claude one evening and untangled it. This article is the result of that conversation, and it’s written for anyone who, like me, wanted to know more about this newfangled world we’ve been thrown in.

Grab a coffee and let’s dive in!

The model does nothing on its own

Here’s the thing that made it click for me: a language model, on its own, is a function. You send it some text, it sends you some text back, and that’s the entire transaction. It can’t run a command, read a file, wait for a result, or remember what it did five minutes ago. Every impressive “agent” you’ve seen doing things on a computer is a plain model wrapped in a lot of software that supplies all of that.

That wrapping is the harness. At its core it’s a loop: the harness sends the model a prompt plus a list of available tools; the model replies with either an answer or “call tool X with these arguments”; the harness actually executes the tool (runs the shell command, reads the file, hits an API), appends the result to the conversation, and hands the whole lot back to the model for the next go. Repeat until the model says it’s done.

Around that loop sit all the boring but essential bits: which tools exist and how they’re described, permission checks before anything destructive runs, the sandbox it all happens in, trimming the conversation when it gets too long, retries when a tool fails, logging, and a way for a human to interrupt. Claude Code is a harness. Codex is a harness. When I chat with Claude in the desktop app, the system prompt, the tool definitions and the thing that executes tool calls are all harness, not model.

If you’re from the 3D world like me, think of it this way: the model is like the character rig, the harness is Unreal Engine. The rig can pose beautifully but does nothing by itself. The engine runs the tick loop, handles input, resolves collisions and lets several actors exist in one level and interact with each other.

A “harness of agents” or “multi-agent harness” is the same idea one level up: a harness that spins up several of these loops, each with its own context and tools, and coordinates them. A lead agent breaks a task into pieces, farms them out to sub-agents, collects the results and decides what happens next. Who talks to whom is harness design, not model behaviour. Two teams with the same model but different harnesses will get wildly different results, which is why so much of this year’s “agent” progress is really progress in harness engineering.

OpenClaw: the open-source original

OpenClaw is an open-source agent harness created by Austrian developer Peter Steinberger, and it’s probably the single project most responsible for pushing the word “harness” into everyday AI chatter this year. Where Claude Code is a harness aimed at coding in a terminal, OpenClaw is aimed at being a personal agent that lives on your machine, talks to you through chat apps, and keeps running when you’re not looking. It can wake up on a schedule, drive a browser, run commands, remember previous sessions and get on with things without you prompting it every step of the way.

You bring your own model: Claude, GPT, Gemini, or a local Qwen served from LM Studio. OpenClaw doesn’t care, it just supplies the loop and everything around it.

It’s had a remarkable year. Jensen Huang called it “the operating system for personal AI” at GTC in March, and Microsoft showed it running natively on Windows at Build in June. That OS analogy is a good one, actually. A harness does for a language model roughly what an OS does for a CPU.

How you install it

It’s not an app in the double-click sense. It’s closer to a background service with a web dashboard and a CLI. You need Node.js (24 recommended, 22.16 or later works) and an API key from a model provider. Then it’s a one-liner:

npx openclaw@latest

The onboarding wizard is smart about existing setups: if it finds a Claude Code or Codex CLI login already on your machine, it verifies it with a real completion, saves the config and opens the dashboard. The core piece is the Gateway, a long-running process on port 18789 that owns the agent loop. Run openclaw gateway install and it registers itself as a LaunchAgent on macOS, a systemd unit on Linux, or a Scheduled Task on Windows, so it survives reboots. There’s a macOS menu bar app that manages the Gateway for you, and a native Windows Hub app too.

Then you bolt channels onto it: WhatsApp, Telegram, Discord, iMessage and others, so you can text your agent from your phone and it acts on your computer. Your customisation lives in ~/.openclaw/openclaw.json (config) and ~/.openclaw/workspace (skills, prompts, memories), and the docs sensibly suggest making the latter a private git repo.

One warning I’ll repeat because it matters: OpenClaw has had critical security vulnerabilities. An always-on agent with tool access and inbound chat channels is a big attack surface. Update promptly, keep the Gateway bound to localhost unless you genuinely need remote access, and read the security guide before pointing it at anything you care about.

Microsoft Autopilot: the enterprise version

On 25 September 2026 Microsoft announced its biggest Copilot update yet, rebuilding the app around three tabs: Home, Code and Autopilot. Autopilot is the interesting one for this article. It was previously called Scout, and Scout is the thing Microsoft demonstrated at Build running on top of OpenClaw. So this isn’t merely a similar idea; it’s the same lineage.

An Autopilot agent gets a name, a role and a goal from its owner, then keeps working inside Microsoft 365 when nobody’s using it. It watches channels, follows up on threads, runs recurring tasks and picks projects back up days later. It lives in your organisation’s tenant with its own identity, memory, computer and workspace, and colleagues can @mention it in Teams or Outlook like a human. It’s cloud-hosted, IT-governed and billed by usage, and it’s entering private preview at the end of September.

Meta Muse: the consumer version

Meta launched Muse on 8 September as a consumer personal-agent app, and it pulled in over half a million users in its first week, hitting number one on the US App Store. The connection to OpenClaw here is unusually explicit. Early adopters noticed that Muse ships with a SOUL.md file, the same Markdown document OpenClaw uses to define an agent’s personality, tone and boundaries, with nearly identical contents.

Someone on X described it as “OpenClaw for normies”, and Nat Friedman, head of product at Meta Superintelligence Labs, didn’t really argue. He confirmed Muse was built from scratch but heavily inspired by OpenClaw, said the team had bought hundreds of Mac minis so staff could use OpenClaw before building their own, and called Steinberger a genius. Steinberger, incidentally, was hired by OpenAI earlier this year, and OpenAI is reportedly working on a response to Muse of its own.

To add to the confusion, there’s also a separate open-source project called OpenMuse from CopilotKit, released in late September as a self-hostable assistant that works with any agent harness. Not Meta’s, and not related to Microsoft either, despite the names.

One pattern, three packagings

  • OpenClaw is the open-source original. You run it yourself, on your own hardware, against whatever model you choose, including local ones. Most work, most control.
  • Autopilot is Microsoft’s enterprise packaging. It lives inside your company’s tenant, governed by IT, and you pay per use. Coming to consumers in 2027.
  • Muse is Meta’s consumer packaging. Download the app, and the agent and your data live in Meta’s cloud. Easiest to try, least control (and Meta gets your data).

What would I actually do with one?

This is the question I kept coming back to. The “it can check multiple apps including chat” bit sounded impressive but abstract, so here are the scenarios that made it concrete for my own workflow.

Overnight render watcher. I kick off a long Blender render or a UE5 cook before bed and message the agent on Discord: watch the output folder, ping me when it’s done or if the log shows errors. It polls the directory, checks the log, and sends me a frame thumbnail when it finishes. If it falls over at 3am, it can restart with lower samples and tell me what it changed.

Comment triage. Both my sites get comments, spam and the occasional real question. A scheduled morning job reads new comments across this site and versluis.com, bins the obvious spam, drafts replies to the real questions in my voice, and sends me one message: three comments worth answering, drafts attached, approve? I reply “1 and 3, skip 2” from my phone and it posts them. Same pattern for YouTube comments.

Research collector. Say I’m planning a video on a new Unreal Engine feature. I tell it once, and it keeps an eye on the Epic forums, the release notes and a couple of subreddits, appending anything relevant to a page in my Notion notes. When I sit down to script the video, the material’s already there.

Cross-app glue. Someone DMs me on Discord asking about a Blender add-on I mentioned in a video. The agent sees the message, searches my blog for the relevant post, checks my GitHub for the repo, and drafts a reply with links for me to approve. I didn’t open any of those three apps.

The pattern in every one of these is the same: something happens in one place, the agent notices, does some legwork across your other tools, and comes to you through whichever chat app is in your pocket with a yes/no decision rather than a task. That “approve?” step is the important bit: I’d want it drafting and asking rather than posting on its own, at least until it’s earned some trust. And given the security history, probably for a good while after that too.

I’ll report back once I’ve had a proper play with OpenClaw on the Mac Studio. If you’ve already got one running, I’d love to hear what you’ve hooked it up to in the comments.

Further reading



If you enjoy my content, please consider supporting me on Ko-fi. In return you can browse this whole site without any pesky ads! More details here.

Leave a Comment!

This site uses Akismet to reduce spam. Learn how your comment data is processed.