5K

One OS User per Agent

Give each long-lived agent its own OS user account and a templated init-system unit, so identity, supervision and privilege come from the host instead of from a per-agent container.

validated in production Updated 2026-09-26 By 5dive contributors (@5dive-bot) Complexity: low Effort: hours Impact: medium
View .md Edit on GitHub
Add to Pack
or

Saved locally in this browser for now.

Cite This Pattern
APA
5dive contributors (@5dive-bot) (2026). One OS User per Agent. In *Awesome Agentic Patterns*. Retrieved September 28, 2026, from https://agentic-patterns.com/patterns/one-os-user-per-agent
BibTeX
@misc{agentic_patterns_one-os-user-per-agent,
  title = {One OS User per Agent},
  author = {5dive contributors (@5dive-bot)},
  year = {2026},
  howpublished = {\url{https://agentic-patterns.com/patterns/one-os-user-per-agent}},
  note = {Awesome Agentic Patterns}
}

Use when

  • Several persistent agents share one host and one working tree
  • You need distinct process identities and per-agent resource accounting
  • Each agent should hold a different privilege scope on the same machine

Avoid when

  • Agents are ephemeral, one per task, and share no state
  • Workloads are untrusted or adversarial and need kernel-level containment
  • The fleet must scale horizontally across many machines

Problem

A fleet of long-lived agents is a different problem from a fleet of ephemeral task runners. The agents stay up for weeks, share one working tree, edit each other's files, and are expected to leave an audit trail a human can read months later.

The reflex is a container or a VM per agent. For this shape it fits badly:

  • The shared working tree fights the boundary. The whole point is that the agents collaborate on one repo, so most of the isolation gets punched back out as bind mounts.
  • You add a second supervisor. The host already has one. Now there is an orchestrator on top of it, with its own restart semantics and its own failure modes.
  • Attribution needs deliberate identity mapping. Containers can provide distinct identities, but shared mounts still need consistent host uid mapping and audit records to identify writers. A separate host uid per persistent agent is one direct way to organize process identity.
  • Per-agent privilege has nowhere to live. Giving one agent the ability to restart a service, and denying it to the others, becomes bespoke policy rather than a line in sudoers.

Meanwhile the host already ships an identity primitive that supports these requirements, and is older than every agent framework on it.

Solution

Make the OS user account the unit of agent identity. One useradd per agent, one templated init unit whose User= is derived mechanically from the instance name, a shared group for the collaboration surface, and per-instance environment overlays for config and credentials.

# one template, N agents: systemctl enable --now myagent@researcher
[Unit]
Description=agent %i
StartLimitIntervalSec=120
StartLimitBurst=10

[Service]
# Identity derives from the instance name; collaboration uses a shared group.
User=agent-%i
Group=shared
# Shared files need group-write permission as well as setgid directories.
UMask=0002
EnvironmentFile=-/etc/agents/common.env
# Per-agent overlay; '-' means an absent file is tolerated.
EnvironmentFile=-/etc/agents/%i.env
ExecStart=/usr/local/bin/agent-start %i
Restart=on-failure
RestartSec=3
# The launcher must use these codes for permanent faults.
RestartPreventExitStatus=2 3

[Install]
WantedBy=multi-user.target

Five things fall out of those two lines:

  1. Process identity is kernel-enforced. ps identifies the uid running a process. File ownership identifies the owner, not the last writer: a shared-group agent can overwrite another agent’s file without changing ownership. Enable OS auditing for writer attribution; Git author metadata is configurable and is not proof of process identity.
  2. Supervision is free. Restart policy, backoff, crash-loop backstop and cgroup accounting come from the init system. RestartPreventExitStatus matters more than it looks: a permanent fault (missing binary, unknown agent type) must fail once and stay failed. Without it, a 3-second restart delay turned one missing binary into 9,563 restarts across a single day.
  3. Collaboration is the group, isolation is the user. The shared work tree is group-owned and setgid. The agents are separate uids in that group. The boundary is drawn once, in the place the kernel already enforces it.
  4. Config and credentials overlay per instance. Keying EnvironmentFile on %i lets two agents of the same type run against different accounts without a second unit.
  5. Privilege is granted per agent, in sudoers. Each agent gets a narrow grant on a mediated CLI, not a broad grant on the filesystem. This is the part most implementations get backwards. See Trade-offs.

Evidence

  • Evidence Grade: medium
  • Most Valuable Findings:
    • Running ~18 persistent agent seats as 18 OS users on a single host, one templated unit, has held in production. Marginal cost of an agent is a uid and a drop-in file.
    • Distinct uids make processes and resource accounting attributable per agent. File ownership remains useful context, but shared writes require audit records to identify the writer.
    • A shared group is a deliberate hole in the boundary, and is where implementations quietly leak (below). Distinct uids do not make shared-group secrets private.
  • Unverified / Unclear: no measurement of how this degrades past a few dozen agents on one host, and no comparison against a rootless-container fleet under the same workload.

How to use it

Reach for this when agents are persistent, collaborating and mutually trusted-ish: an internal fleet on your own hardware, not untrusted tenant code.

  1. Create one system user per agent, all in a shared group. Home directory per user.
  2. Make the shared work tree group-owned and setgid, and set UMask=0002 so new shared files are group-writable. Keep secrets outside it, with explicitly restrictive file and directory modes.
  3. Write one templated unit with User=agent-%i, per-instance EnvironmentFile overlays, and a restart policy that distinguishes transient from permanent failure.
  4. Keep credentials in private per-user directories or root-owned service configuration. Environment-injected secrets remain accessible to that agent. For elevated capabilities, grant only a narrow mediated command through sudoers; inspect its arguments and effects.
  5. Enable an instance per agent. Adding agent number nineteen is useradd plus systemctl enable --now.

Trade-offs

Pros

  • No image to build, ship or rebuild. The agent runs against the host's real toolchain.
  • Supervision, restart, log capture and cgroup limits come from the init system already running.
  • Per-agent process identity and accounting; OS auditing can record writes to shared files.
  • Per-agent privilege is a sudoers line, reviewable in one file.
  • Survives reboot, and the marginal cost per agent is a uid.

Cons

  • A uid is not a sandbox. Same kernel, same network namespace, no syscall filtering unless you add it. Against an adversarial agent this is the wrong tool. Use a VM.
  • The shared group leaks by design. Anything group-readable is readable by every agent in the group. Private uid-owned directories can isolate credentials from other uids when permissions are restrictive; shared-group files cannot. Root-owned configuration and mediated commands are options for higher-privilege secrets, not guarantees that an agent cannot read its own tokens.
  • Root compromise is fleet-wide. There is one kernel and one root.
  • Host-bound. Scaling past one machine needs machinery this pattern does not provide.
  • Cross-user writes need the whole chain composed. Group membership is necessary and not sufficient: a directory's setgid bit, the parent directories' traversal bits, and hardening sysctls such as fs.protected_regular each independently block a write that "should" work. Expect to debug a permission denial that looks like a missing grant and is not.

References