A security sandbox
the boundaryKeeps a prompt-injected or simply wrong agent off the rest of your machine: your other projects, your SSH keys, your LAN, your shell startup files.
sandy · an isolated sibling for your coding agents
sandy runs Claude Code, Gemini CLI, Codex, OpenCode and Grok in a per-project Docker sandbox with their permission prompts switched off. The agent works without stopping to ask, and the operating system, not the agent, decides what it can touch. Press Next to see what that means, watch the boundary turn things away, and get it running.
The invariantThe sandbox's guarantees do not depend on the agent behaving. With the default settings, an agent that has been prompt-injected, and holds every permission it asks for, still cannot plant a git hook, reach your home network, or read a credential it was not given.
~/dev/myapp rw.git/hooks ro.github/workflows ro~/.ssh not mounted192.168.0.0/16 no route
What it is
Keeps a prompt-injected or simply wrong agent off the rest of your machine: your other projects, your SSH keys, your LAN, your shell startup files.
Each project gets its own agent home: plugins, memory, history, credentials and installed packages. Nothing bleeds between projects. This half is useful even if you trust the agent completely.
~/.sandy/sandboxes/ ├── webapp-a1b2c3d4/ # one sandbox per project, named from its path │ ├── claude/ # mounted at ~/.claude: plugins, memory, history │ ├── codex/ # and one home per agent you use there │ ├── pip/ npm-global/ go/ cargo/ uv/ # packages you install persist │ └── WORKSPACE.json └── data-pipeline-e5f6a7b8/ └── claude/ # a different project, a different agent home
The problem
It reads files it did not write, and any of them can carry instructions. Give it autonomy and it acts on them with your user's full reach. Three kinds of reach matter.
Everything your user can write: other repositories, your shell startup files, git hooks that run on your next checkout, the CI workflow that runs on your next push.
Your home or office LAN, the router's admin page, a NAS, services on localhost, and cloud metadata endpoints that hand out credentials.
Your SSH keys, your cloud configs, and the agent's own long-lived refresh token. Anything it can read, it can send somewhere.
The problem
Asking before every command protects you only if you read every command. After the fortieth prompt nobody does. sandy moves the decision out of the agent and into the operating system.
Prompts are off, so it reads, edits, builds and tests without waiting on you.
Mounts, network namespaces and dropped capabilities refuse what the sandbox does not allow, however the request was phrased.
The work lands in your repository, where you review it in git. sandy tells you at session end about anything that needs a look.
The layers
One sandbox per workspace, one sandy per workspace at a time. Projects never share agent state.
~/.sandy/sandboxes/<name>-<hash>Read-only root filesystem, a throwaway home, and only the project directory mounted from the host. About 25 sensitive paths inside it are read-only too.
--read-only · -v …:roThe agent sits on a network with no route out except a small proxy. Private and LAN addresses are refused; non-TCP traffic is dropped by the topology.
--internal · SNI proxyEach agent gets only its own login. Claude's is read fresh from your host at every launch and never stored in the sandbox. No SSH private key enters unless you name it.
SANDY_SSH_KEYS · SANDY_SUSPICIOUSNon-root, every capability dropped, no new privileges, Docker's default seccomp and AppArmor profiles, and no Docker socket.
--cap-drop ALL · no-new-privilegesCPU, memory, process count and temporary storage are capped, sized from your machine, so a runaway build can't take the host down with it.
--cpus · --memory · --pids-limitThe layers · network
The proxy works the same on macOS and Linux because the isolation comes from Docker's network layout, not from firewall rules on the host. Most people never change the default.
| SANDY_EGRESS | The agent can reach | From a repo's committed config |
|---|---|---|
| permissive default | The public internet. Not your LAN, not the host, not cloud metadata, not well-known DNS-over-HTTPS resolvers. | asks you It is looser than strict, and a repo's value outranks yours. |
| strict | Only an allowlist: the model providers, GitHub, the npm, PyPI, crates, Go and Debian registries, and hosts you add with SANDY_ALLOW_HOSTS. | free It only tightens. |
| off | Linux: public internet, LAN blocked by iptables. macOS: everything, including your LAN. | asks you It loosens. |
The proxy reads the hostname from the TLS handshake and never decrypts. It logs what it refused, and with SANDY_EGRESS_LOG=1 every distinct host the session reached, so after a suspicious session you can answer "what did it talk to?"
The seam
This is everything the agent can see of your machine. Anything not in the table is not mounted at all, so there is nothing for the agent to read, however it asks.
| In the container | Comes from | Mode | Holds |
|---|---|---|---|
| / | the image | read-only | Python 3.13, Node 24, Go 1.26, Rust, C/C++, git, uv and the agent CLIs. |
| ~/dev/myapp | your project | read-write | Your code. The one host directory the agent edits. |
| .git/hooks/ .git/config | your project | read-only | About 25 protected paths: git hooks and config, shell startup files, .github/workflows/, .vscode/, .sandy/, and more. |
| ~/.claude, ~/.codex, … | this project's sandbox | read-write | Plugins, memory and history for this project only. |
| ~/.pip-packages, ~/go, … | this project's sandbox | read-write | Packages you install. They survive relaunches. |
| ~/.claude/.credentials.json | your host, read each launch | ephemeral | The agent's login. Never written into the sandbox. |
| /etc/sandy-session.json | sandy | read-only | What this session is: version, network posture, credential mode. |
| ~/.ssh, ~/.aws, other repos | — | not mounted | Nothing to read. Named SSH keys are the one opt-in. |
| /var/run/docker.sock | — | not mounted | The agent cannot start containers or reach the Docker daemon. |
Trust tiers
A project can commit a .sandy/config so everyone who clones it gets the same setup. Clone someone else's repository and that file is theirs, so sandy treats it accordingly. It is read as KEY=VALUE lines and never run as a script.
SANDY_AGENTSANDY_MODELSANDY_EFFORTSANDY_EGRESS=strictSANDY_SUSPICIOUS=1SANDY_CPUSSANDY_MEMSANDY_SKILL_PACKSSANDY_EGRESS=offSANDY_EGRESS=permissiveSANDY_SSHSANDY_SSH_KEYSSANDY_ALLOW_HOSTSSANDY_EXTRA_ENVSANDY_AGENT_ARGS*_API_KEYSettings in your own ~/.sandy/config never ask: that file is yours. Approval is per workspace and remembered, and any change to the approved keys asks again.
Trust tiers
.sandy/Dockerfile is new or changed. You see all of it before it builds.When nobody can answer (a one-shot sandy -p, a background start with no terminal), every gate fails closed. sandy --approvals reports what each gate would decide for a workspace without granting anything.
Knowing it holds
/etc/sandy-session.json says the session is sandy, which version, which network posture and which kind of credential is present. It is mounted read-only, so a repository cannot forge it. In-container tools should trust it over guesses from uid or environment variables.
sandy --print-state lists every sandbox, image and running session as one JSON document, and sandy --doctor checks the host and flags stale locks, orphaned networks and outdated images.
A protected file that appeared during the session, a git branch left switched, a permission mode the agent changed, an image refresh that had to wait. Each gets a line when the session ends.
Inside a session, curl -m 5 http://192.168.1.1 should fail and curl -m 5 https://api.anthropic.com should succeed.
Knowing it holds
A sandbox you can trust is one whose limits are written down. These are the main ones; the threat model has the rest.
SANDY_EGRESS=off on macOS leaves the LAN reachable. Keep the default proxy on.SANDY_SUSPICIOUS=1 strips the refresh token and defaults to strict egress.Knowing it holds
Other good tools attack the same problem from different angles. Each picks a different trade, as of October 2026.
| sandy | Docker Sandboxes | NVIDIA OpenShell | Built-in agent sandbox | |
|---|---|---|---|---|
| Boundary | Hardened container, shared kernel | microVM, its own kernel | Container or VM, plus Landlock and seccomp | Shell commands only |
| You install | One script, on the Docker you have | Its CLI and VM runtime | A gateway and supervisor | Nothing |
| Agents | Five, up to four side by side | Eleven | Several, via providers | Itself |
| Per-project agent state | Yes | Per sandbox | Per sandbox | No, global |
| Credentials | Mounted per session, never stored | Injected by a proxy, never inside | Injected by a proxy, never seen | The agent's own |
| Network | Host allowlist proxy, no TLS decryption | Default deny, host and method/path rules | Default deny, L7 rules | Domain rules for shell commands |
| A repo you didn't write | Hooks, CI and shell rc read-only; approval gates | Clone mode, or review changes | Filesystem policy | Partial |
| Cost | Free, MIT | Free locally; paid governance | Apache 2.0, early | Free |
Choose a microVM for a hypervisor boundary, Windows, or Docker inside the sandbox. Choose sandy for your existing Docker, repos you didn't write, several agents, and separate per-project setups. The README's Why sandy section has the full comparison and its sources.
Many agents
Every agent gets the same boundary and its own home directory and credentials inside the project's sandbox. Pick one per project, or run up to four at once.
Pro or Max login from your host, an API key, or a Console profile.
An API key, your Google login, or gcloud credentials.
An API key, or a ChatGPT login that persists per project.
Any provider, including a local model on your LAN through one narrow opening.
An xAI API key, or an interactive login.
SANDY_AGENT=claude,codex # two panes, side by side SANDY_AGENT=all # claude, gemini, codex, opencode
One tmux session, one shared workspace, separate agent homes. Closing one pane leaves the others running.
Run it · 1 of 4
curl -fsSL https://raw.githubusercontent.com/rappdw/sandy/main/doctor.sh | bash
The doctor checks that Docker is reachable, that git and curl are present, that your agent credentials can be found and that ~/.local/bin is on your PATH. For anything missing it prints a command to copy.
Run it · 2 of 4
curl -fsSL https://raw.githubusercontent.com/rappdw/sandy/main/install.sh | bash
cd ~/dev/myapp sandy # interactive session sandy -p "summarise src/" # or a one-shot prompt
The installer puts one script in ~/.local/bin. The first launch builds the images, which takes a while once; after that, sessions start in seconds and the images rebuild themselves when your agent releases an update. Claude Pro and Max logins are picked up from your host, so there is no API key to set.
Run it · 3 of 4
SANDY_AGENT=claude
SANDY_EFFORT=high
SANDY_EGRESS=strict # tighter than the default
SANDY_MEM=8gPrecedence runs from a command-line flag, to your shell environment, to the project file, to your own defaults.
Run it · 4 of 4
sandy --start # start detached; survives a closed terminal or a reboot sandy --attach # attach from any terminal or over ssh sandy --stop # tear it down
sandy --doctor # host and runtime health sandy --gc --dry-run # show unused images and networks sandy --reset-sandbox # rebuild a sandbox you no longer trust sandy --upgrade # update sandy itself
Daemon mode is what sandy-ui uses to keep a session alive across VS Code restarts. Every maintenance command prints its plan first, and --dry-run stops there.
sandy
sandy is one MIT-licensed bash script. It runs on any Docker-compatible runtime on macOS or Linux and needs no server or account of its own.