Quick Takeaway: Operating an autonomous OpenClaw agent stack 24/7 in production demands a dedicated KVM VPS provisioned with at least 4 vCPUs, 8 GB RAM (preventing OOM panics during concurrent headless browser execution), 80 GB PCIe 4.0 NVMe storage, and an unthrottled 1 Gbps uplink running TCP BBR to minimize p99 latency across continuous upstream LLM API roundtrips. To ensure deterministic stability, agent worker pools must be isolated via cgroups v2 with strict memory.high and CPUQuota limits alongside kernel tuning (vm.max_map_count = 262144, net.ipv4.tcp_fastopen = 3, and fs.file-max = 2097152) to eliminate socket exhaustion during recursive tool invocation. Production deployment further mandates a sandboxed systemd service (ProtectSystem=strict, NoNewPrivileges=true) coupled with an automated watchdog to guarantee zero-downtime recovery without state corruption.
Table of Contents
- Hardware Sizing and System Requirements for 24/7 Autonomous Agent Execution
- Isolating the Execution Environment: Docker Sandboxing and cgroups v2
- Deploying OpenClaw with Docker Compose and API Provider Bridges
- Continuous Task Scheduling, Headless Browser Drivers, and Webhooks
- API Key Security, Rate Limiting, and Cost Containment Strategies
- Systemd Process Watchdogs, Monitoring, and Recovery from Runaway Loops
- Frequently Asked Questions (FAQ)
Hardware Sizing and System Requirements for 24/7 Autonomous Agent Execution
Running an autonomous execution framework like openclaw on VPS 24/7 shifts compute demands away from traditional request-response web models toward sustained, asynchronous event loops, headless browser automation, and concurrent local state persistence. An autonomous agent does not idle between user requests; it cycles through perception-action loops, dispatches multi-step tool calls, performs DOM inspection via the Chrome DevTools Protocol (CDP), and reads or writes continuous vector embeddings.
Under-provisioned infrastructure leads directly to memory leaks, cascading process crashes, and event loop starvation. Deploying openclaw on vps 24/7 requires matching hardware resources directly against concurrent execution workers, headless browser sandboxes, and disk I/O characteristics.
Compute Architecture: Single-Thread Frequency vs. Multi-Core Scaling
Autonomous agent workloads exhibit a dual compute profile: 1. Single-threaded, latency-critical orchestration: The primary agent process (typically executing on Node.js or Python) coordinates state machines, parses structured JSON outputs, compiles prompt templates, and handles WebSocket streaming. This orchestration layer is bound to a single thread; its speed depends directly on the core’s base clock speed and instructions per clock (IPC), making high-frequency architectures (such as AMD EPYC 9004/7003 series or AMD Ryzen 9 exceeding 3.7–4.5 GHz) vastly superior to legacy multi-socket Xeons with low per-core clocks. 2. Multi-process child tasking: When tools trigger sandboxed Python evaluation, bash scripts, or Playwright/Puppeteer browser contexts, the operating system forks separate child processes. Each active headless browser instance requires dedicated compute to avoid throttling page rendering and JavaScript execution.
The Operational Cost of CPU Steal Time (%st)
In shared or oversubscribed cloud environments, hypervisors allocate physical CPU cycles across multiple noisy tenants. When an autonomous agent executes a multi-step verification sequence, CPU steal time (%st) introduces unpredictable execution delays.
Hypervisor Scheduling Delay
┌─────────────────────────┐
Primary Agent Thread │ │ Context Resumed
──────────────────────────────►│ CPU STEAL (%st > 3%) ├────────────────────────►
(Active CDP / WebSocket Link) │ │ Socket Keepalive Missed
└─────────────────────────┘ ──► ETIMEDOUT / TCP RST
If an agent thread is paused for even 300–500 ms while holding an active CDP session or an HTTP chunked transfer connection to an inference endpoint, the socket may drop, causing unhandled promise rejections, lost session context, and corrupted state machines.
To audit CPU steal on your instance, monitor core metrics in real time:
# Sample CPU metrics across all cores at 1-second intervals for 5 iterations
mpstat -P ALL 1 5
Inspect the %st column. For resilient 24/7 agent operations, %st must remain strictly at 0.0%. Deploying on dedicated KVM virtualization—such as the bare-metal sliced KVM nodes on tropic.host—guarantees zero hypervisor-level CPU oversubscription, providing dedicated physical core pinning without noisy-neighbor contention.
Memory Allocation and the Mechanics of Headless Browser Leaks
Memory consumption in 24/7 agent pipelines is dominated by three discrete layers:
- Agent Supervisor Runtime: The core runtime (Node.js/Python) consumes an initial 200–450 MB of Resident Set Size (RSS). Over days of continuous runtime, heap fragmentation can expand this footprint to 800 MB–1.2 GB unless garbage collection is aggressively managed.
- Local Vector Caching and State Management: If the agent uses embedded vector engines (such as Chroma, LanceDB, or SQLite with
sqlite-vec) for short-term memory retrieval, memory-mapped files (mmap) require sufficient page cache space to prevent constant disk reads. Allocate a minimum of 1 GB to 2 GB solely for in-memory index caching. - Headless Browser Contexts (Chromium / WebKit): A single Chromium instance spawns multiple helper processes: the browser broker, GPU process, network service, and separate renderer processes for each open tab or iframe.
┌────────────────────────────────────────────────────────┐
│ Chromium Process Tree (Per Session) │
├───────────────────────────────┬────────────────────────┤
│ Process Role │ Resident Memory (RSS) │
├───────────────────────────────┼────────────────────────┤
│ Browser Main (Broker) │ 95 MB – 140 MB │
│ Network Service │ 45 MB – 80 MB │
│ GPU Process │ 60 MB – 110 MB │
│ Renderer Tab 1 (Target Page) │ 180 MB – 450 MB │
│ Dedicated DevTools Session │ 40 MB – 75 MB │
├───────────────────────────────┼────────────────────────┤
│ Total Per Isolated Context │ 420 MB – 855 MB │
└───────────────────────────────┴────────────────────────┘
If an agent workflow opens 4 concurrent worker threads to scrape or execute transactions, the browser subsystem alone demands 2.5 GB to 3.5 GB of available RAM. Without strict cgroups limits, a single JavaScript-heavy single-page application (SPA) can trigger a memory leak that forces the Linux Out-Of-Memory (OOM) killer to terminate the parent supervisor process.
Disk I/O Profiles: WAL Contention and SQLite Concurrency
Autonomous agents continuously persist telemetry, session checkpoints, scratchpads, and execution graphs. Most production configurations use an embedded SQLite database configured in Write-Ahead Logging (WAL) mode or local file-backed document stores.
In WAL mode, writes append sequentially to the -wal file while reads query both the main database and the log. When the WAL file reaches its threshold (typically 1,000 pages), SQLite attempts a checkpoint (PRAGMA wal_checkpoint(PASSIVE) or TRUNCATE). On low-tier cloud storage with high latency or constrained IOPS, this checkpoint locks the database file. If an agent tries to log an execution step during an I/O freeze, the operation raises an SQLITE_BUSY error, halting the loop.
Enterprise PCIe 4.0 NVMe drives—such as those provisioned across tropic.host cloud infrastructure—deliver random 4K QD1 read/write performance exceeding 50,000 IOPS. This prevents WAL write locks and ensures low p99 latency during continuous logging.
To verify your storage subsystem under realistic SQLite concurrency profiles, execute a random read/write fio benchmark:
fio --name=wal_simulation \
--ioengine=libaio \
--iodepth=1 \
--rw=randrw \
--rwmixwrite=70 \
--bs=4k \
--direct=1 \
--size=1G \
--numjobs=4 \
--runtime=30 \
--group_reporting
Ensure that write latency at the 99th percentile (clat p99) remains below 1.2 ms under load.
Hardware Sizing and Concurrency Matrix
The following hardware matrix outlines minimum and recommended configurations for running openclaw on vps 24/7 across three distinct workload tiers.
| Operational Tier | Target Workload Profile | vCPU Spec (Dedicated KVM) | RAM Allocation | Storage Tier & Throughput | Max Concurrent Agents / Browsers | Target Latency (Tool Execution p99) | Recommended Infrastructure Profile |
|---|---|---|---|---|---|---|---|
| Tier 1: Minimal | Single-agent execution, text-only APIs, webhooks, lightweight scraping (no Chromium). | 2 vCPU (AMD EPYC/Ryzen @ 3.2+ GHz) | 4 GB ECC DDR4/DDR5 | 40 GB NVMe (Random 4K > 15k IOPS) | 1 Active Loop / 0 Browsers (cURL/Cheerio only) | < 850 ms | tropic.host Starter KVM |
| Tier 2: Production | Continuous multi-agent routing, 1–3 concurrent Playwright browser sessions, local Chroma/SQLite storage. | 4 vCPU (High-IPC AMD @ 3.7+ GHz) | 8 GB – 16 GB DDR5 | 80–120 GB PCIe 4.0 NVMe (Random 4K > 40k IOPS) | 4 Active Loops / 3 Headless Browser Tabs | < 320 ms | tropic.host Pro KVM |
| Tier 3: Enterprise Swarm | Swarm coordination, parallel browser automation (10+ contexts), continuous embeddings, local OCR/Whisper. | 8–16 vCPU (Dedicated AMD EPYC / Ryzen 9) | 32 GB – 64 GB DDR5 | 250+ GB Enterprise NVMe (Random 4K > 65k IOPS) | 16 Active Loops / 12+ Concurrent Browsers | < 110 ms | tropic.host Enterprise NVMe |
Production Kernel and OS Tuning (sysctl.conf)
The standard Linux kernel network stack and virtual memory parameters are tuned for batch-oriented desktop or standard web server workloads. Sustaining thousands of long-lived WebSocket connections, rapid loopback inter-process communications, and intensive SQLite file memory mapping requires specific kernel adjustments.
Deploy the following configuration to /etc/sysctl.d/99-openclaw.conf:
# Virtual Memory Management
# Lower swappiness to prevent swapping working process memory to disk
vm.swappiness = 10
# Increase maximum memory-mapped allocations for vector databases (Chroma/Qdrant)
vm.max_map_count = 262144
# Keep dirty pages low to trigger continuous background writeback rather than large disk flushes
vm.dirty_background_ratio = 5
vm.dirty_ratio = 15
# Network Core Tuning
# Increase the queue length for incoming connections to avoid dropped syn packets during burst spikes
net.core.somaxconn = 4096
net.ipv4.tcp_max_syn_backlog = 4096
# Allow immediate reuse of sockets in TIME_WAIT state for outgoing API connections
net.ipv4.tcp_tw_reuse = 1
# Modern TCP congestion control: enforce BBR (requires Linux kernel >= 4.9)
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr
# Increase system file descriptors to handle simultaneous sockets, temp files, and pipes
fs.file-max = 2097152
fs.inotify.max_user_watches = 524288
Apply these settings without rebooting:
sysctl --system
Verify that the BBR algorithm is active on your network interface:
sysctl net.ipv4.tcp_congestion_control
# Output must read: net.ipv4.tcp_congestion_control = bbr
Process Isolation and Cgroups v2 Protection
To prevent rogue child processes (e.g., a stalled Chromium instance or an infinite loop in a dynamically executed Python script) from exhausting node memory and killing the agent supervisor, isolate the agent using a dedicated systemd service backed by cgroups v2.
Create the service unit file at /etc/systemd/system/openclaw.service:
[Unit]
Description=OpenClaw Autonomous Agent Supervisor
After=network.target network-online.target
Wants=network-online.target
[Service]
Type=simple
User=openclaw
Group=openclaw
WorkingDirectory=/opt/openclaw
Environment="NODE_ENV=production"
ExecStart=/usr/bin/node /opt/openclaw/dist/server.js
Restart=always
RestartSec=5s
# Cgroups v2 Resource Accounting & Limits
CPUAccounting=true
CPUWeight=100
MemoryAccounting=true
# Set memory throttling threshold before hard OOM termination
MemoryHigh=12G
# Strict ceiling: triggers cgroups-level killer inside the unit, not host-wide
MemoryMax=14G
# Protect agent supervisor from the system OOM killer
# Negative values prioritize this process during memory pressure (-1000 to 1000)
OOMScoreAdjust=-500
# File descriptor ceiling
LimitNOFILE=65535
LimitNPROC=4096
# Security Sandboxing
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/opt/openclaw/data /opt/openclaw/logs /tmp
[Install]
WantedBy=multi-user.target
Reload the systemd daemon, enable, and start the isolated service:
systemctl daemon-reload
systemctl enable --now openclaw.service
Verify that cgroups v2 limits are active by checking the controller hierarchy:
systemd-cgls /system.slice/openclaw.service
This configuration ensures that even if child processes spike RAM usage under heavy load, the supervisor stays running, logs the bottleneck, and gracefully terminates only the misbehaving task. Pairing these operating system safeguards with clean, high-performance hardware from tropic.host—equipped with dedicated KVM virtualization, zero CPU steal (%st = 0.0%), and enterprise PCIe 4.0 NVMe drives—provides the exact hardware foundation required for deterministic 24/7 autonomous agent uptime.
Isolating the Execution Environment: Docker Sandboxing and cgroups v2
Host-level isolation via systemd secures the OpenClaw orchestration process, but running arbitrary bash commands, Python scripts, and headless browser sessions requested by LLM tool calls directly within that process boundary creates an unacceptable attack surface. When deploying openclaw on vps 24/7, runtime tool execution must be decoupled from the supervisor and confined to ephemeral, least-privilege OCI containers. Without strict kernel-level containment, a single prompt injection or hallucinated destructive command (rm -rf /*, unconstrained fork bombs, or lateral network scans) can compromise the host, leak credentials, or trigger null-routing from your hosting provider due to outbound abuse traffic.
True isolation requires a multi-layered defense: configuring the Docker daemon to leverage the unified cgroups v2 hierarchy, stripping Linux capabilities, enforcing immutable root filesystems, restricting inter-container communications, and applying deterministic compute ceilings.
┌───────────────────────────────────────────────┐
│ tropic.host KVM VPS │
│ Host OS (Ubuntu 24.04 LTS / Debian 12)│
│ sysctl, cgroups v2 unified tree │
└───────────────────────┬───────────────────────┘
│
cgroups v2 Enforcement │ systemd cgroupdriver
▼
┌───────────────────────────────────────────────┐
│ systemd Slice: openclaw.service │
│ (Node.js Orchestrator & Task Scheduler) │
└───────────────────────┬───────────────────────┘
│ Spawns isolated execution
│ via Docker Socket / API
▼
┌───────────────────────────────────────────────┐
│ Isolated OCI Sandbox (openclaw-runner) │
│ ┌─────────────────────────────────────────┐ │
│ │ Security: cap_drop: [ALL], read_only │ │
│ │ Memory: max 2G, high 1.7G, swap = 0 │ │
│ │ CPU: cpus: 2.0 (CFS: 200000/100000) │ │
│ │ PIDs: pids.max = 128 (Anti-forkbomb) │ │
│ │ Net: DOCKER-USER egress filter (No SSRF)│ │
│ └─────────────────────────────────────────┘ │
└───────────────────────────────────────────────┘
Step 1: Enforce the cgroups v2 Driver in the Docker Daemon
Modern Linux distributions mount cgroups v2 by default under /sys/fs/cgroup. Verify that the host is operating in full unified mode before provisioning container runtimes:
[ -f /sys/fs/cgroup/cgroup.controllers ] && echo "cgroups v2 active" || echo "Legacy cgroups v1 detected"
If the host reports legacy v1, append systemd.unified_cgroup_hierarchy=1 to GRUB_CMDLINE_LINUX in /etc/default/grub, run update-grub, and reboot.
Once unified hierarchy is verified, configure /etc/docker/daemon.json to enforce the systemd cgroup driver, disable inter-container communication on default bridges, eliminate userland proxy overhead, and lock down process privilege escalation:
{
"exec-opts": ["native.cgroupdriver=systemd"],
"log-driver": "json-file",
"log-opts": {
"max-size": "50m",
"max-file": "3"
},
"live-restore": true,
"userland-proxy": false,
"icc": false,
"no-new-privileges": true,
"default-ulimits": {
"nofile": {
"Name": "nofile",
"Hard": 65535,
"Soft": 65535
},
"nproc": {
"Name": "nproc",
"Hard": 1024,
"Soft": 512
}
}
}
Applying "icc": false ensures that an untrusted agent execution container cannot sniff or interact with network packets destined for other co-located containers (such as your Postgres metadata database or Redis cache). Restart the Docker daemon to apply:
systemctl restart docker
docker info | grep -i cgroup
The output must show Cgroup Driver: systemd and Cgroup Version: 2.
Step 2: Deploying the Hardened Ephemeral Execution Sandbox
OpenClaw executes dynamic actions via an isolated runner image (openclaw-runner:latest). Rather than granting the runner full root access to the filesystem, the container must be instantiated with dropped kernel capabilities, an immutable root filesystem (read_only: true), and strictly bounded ephemeral mounts.
Below is the production-hardened docker-compose.sandbox.yml specification used to orchestrate runner environments:
services:
agent-sandbox:
image: openclaw-runner:latest
container_name: openclaw-sandbox-ephemeral
restart: "no"
user: "10001:10001"
read_only: true
security_opt:
- no-new-privileges:true
- seccomp=/etc/openclaw/seccomp-profile.json
cap_drop:
- ALL
cap_add:
- CHOWN
- SETUID
- SETGID
deploy:
resources:
limits:
cpus: "2.00"
memory: 2048M
pids: 128
reservations:
cpus: "0.50"
memory: 512M
tmpfs:
- /tmp:rw,noexec,nosuid,nodev,size=256m
- /run:rw,noexec,nosuid,nodev,size=64m
- /home/sandbox:rw,exec,nosuid,nodev,size=1024m
networks:
- openclaw-isolated-sandbox
dns:
- 1.1.1.1
- 8.8.8.8
environment:
- NODE_ENV=production
- TMPDIR=/tmp
- HOME=/home/sandbox
networks:
openclaw-isolated-sandbox:
driver: bridge
internal: false
driver_opts:
com.docker.network.bridge.enable_icc: "false"
com.docker.network.bridge.enable_ip_masquerade: "true"
Breakdown of Security Directives:
read_only: true: Blocks runtime modification of system binaries (/bin,/usr,/lib). Any attempt by an LLM-spawned payload to install persistent rootkits, modify Python site-packages, or overwrite dynamic libraries fails instantly withRead-only file system.cap_drop: [ALL]: Strips all 41 Linux capabilities. The sandboxed code cannot interact with kernel network routing (CAP_NET_ADMIN), mount filesystems (CAP_SYS_ADMIN), trace other processes (CAP_SYS_PTRACE), or override DAC permissions (CAP_DAC_OVERRIDE).tmpfs with noexec,nosuid,nodev: Memory-backed storage partitions for/tmpand/runensure fast temporary file generation without disk wear on the host. Mounting withnoexecprevents an agent from downloading a raw binary (e.g., viacurl) into/tmpand executing it directly. The dedicated working directory/home/sandboxallows execution exclusively for interpreted code files generated during valid tool runs.pids: 128: Imposes a strict process ceiling through the cgroups v2pids.maxcontroller. If an agent script executes a fork bomb (:(){ :|:& };:), process generation halts abruptly once the count exceeds 128, preventing host thread exhaustion and kernel panics.
Step 3: Granular cgroups v2 Resource Accounting
Under cgroups v2, resource tracking is unified under /sys/fs/cgroup/system.slice/docker-<container_id>.scope/. When running OpenClaw tasks, the engine dynamically sets CFS (Completely Fair Scheduler) bandwidth constraints and multi-tiered memory thresholds.
To verify how the limits defined in your compose file translate directly into kernel controls, inspect the live container cgroups parameters while an agent task executes:
# Retrieve container ID
CONTAINER_ID=$(docker ps -aqf "name=openclaw-sandbox-ephemeral")
# Inspect CPU quotas (CFS bandwidth control)
cat /sys/fs/cgroup/system.slice/docker-${CONTAINER_ID}.scope/cpu.max
# Output: 200000 100000 (Allows 200,000us of runtime per 100,000us period = 2 vCPUs)
# Inspect Memory Ceiling and Throttling Limits
cat /sys/fs/cgroup/system.slice/docker-${CONTAINER_ID}.scope/memory.max
# Output: 2147483648 (Exact 2048M hard ceiling)
cat /sys/fs/cgroup/system.slice/docker-${CONTAINER_ID}.scope/memory.swap.max
# Output: 0 (Swap completely disabled for the container to avoid disk thrashing)
Setting memory.swap.max to 0 is critical for real-time deterministic performance. On unmanaged environments, when an agent processes massive context windows or runs out-of-memory computations, the kernel attempts to swap anonymous memory pages to disk, causing storage queue depths to spike and pushing I/O latency past acceptable p99 thresholds.
Deploying openclaw on vps 24/7 on tropic.host ensures this boundary is enforced reliably: dedicated KVM hypervisor instances with zero CPU steal (%st = 0.0%) and enterprise PCIe 4.0 NVMe drives prevent host-level I/O stalls even when agent containers hit memory ceilings and trigger instant, isolated process terminations.
Step 4: Mitigating SSRF and Enforcing Egress Boundaries via iptables
An autonomous agent with open Internet access can inadvertently be weaponized as an internal network proxy. If an attacker injects a prompt instructing OpenClaw to fetch http://192.168.1.1 or query local host services, the agent must be blocked at the network packet filtering layer.
Docker bypasses standard UFW rules by inserting its own chains into iptables. To secure container traffic, insert custom drops into the DOCKER-USER chain to prevent containers on the sandbox subnet from reaching host loopback addresses, local private subnets (RFC 1918), and link-local interfaces:
# Identify your sandbox subnet (e.g., 172.28.0.0/16)
SANDBOX_SUBNET="172.28.0.0/16"
# Allow established and related connections
iptables -I DOCKER-USER -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT
# Block egress from sandbox to RFC 1918 private IPv4 networks (SSRF prevention)
iptables -A DOCKER-USER -s ${SANDBOX_SUBNET} -d 10.0.0.0/8 -j DROP
iptables -A DOCKER-USER -s ${SANDBOX_SUBNET} -d 172.16.0.0/12 -j DROP
iptables -A DOCKER-USER -s ${SANDBOX_SUBNET} -d 192.168.0.0/16 -j DROP
# Block egress to link-local and cloud metadata endpoints
iptables -A DOCKER-USER -s ${SANDBOX_SUBNET} -d 169.254.169.254/32 -j DROP
# Block direct access to host docker daemon socket and local loopback range
iptables -A DOCKER-USER -s ${SANDBOX_SUBNET} -d 127.0.0.0/8 -j DROP
# Persist iptables across reboots
netfilter-persistent save
For high-throughput tool runs (e.g., processing thousands of external web pages per hour), tune the host network stack in /etc/sysctl.d/99-openclaw-sandbox.conf to handle rapid socket recycling without exhausting ephemeral ports:
# Accelerate socket recycling and expand ephemeral port range
net.ipv4.ip_local_port_range = 10240 65535
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
# Protect host from syn flood during untrusted outbound handshakes
net.ipv4.tcp_syncookies = 1
net.ipv4.tcp_max_syn_backlog = 8192
net.core.somaxconn = 4096
Reload the sysctl configuration:
sysctl --system
Step 5: Validating the Sandbox Security Perimeter
Before handing autonomous execution privileges to OpenClaw, run automated validation tests against the sandbox to prove that privilege escalation, network escaping, and resource exhaustion are completely contained.
Test 1: Fork Bomb and PID Controller Ceiling
Run an aggressive fork loop inside the sandbox container to verify that pids.max halts process generation without affecting the host:
docker run --rm \
--pids-limit 128 \
--cap-drop ALL \
openclaw-runner:latest \
bash -c ':(){ :|:& };:'
Expected Result: The bash terminal outputs bash: fork: retry: Resource temporarily unavailable and execution halts safely when PID count hits 128. The host OS remains completely unaffected, showing 0.0% CPU lockup.
Test 2: Enforced Read-Only Filesystem Verification
Verify that root filesystem modification is blocked:
docker run --rm \
--read-only \
--tmpfs /tmp:rw,noexec,size=64m \
openclaw-runner:latest \
touch /usr/local/bin/malicious_tool
Expected Result: touch: cannot touch '/usr/local/bin/malicious_tool': Read-only file system.
Test 3: SSRF Network Egress Filtering
Verify that the DOCKER-USER iptables rules drop unauthorized internal network traversal:
docker run --rm \
--network openclaw-isolated-sandbox \
openclaw-runner:latest \
curl --connect-timeout 3 -s http://192.168.1.1
Expected Result: The connection hangs until timeout and outputs curl: (28) Failed to connect to 192.168.1.1 port 80: Connection timed out, while public internet calls (curl -s https://api.github.com) resolve instantly.
By locking the execution environment inside a capability-stripped OCI container backed by cgroups v2 resource ceilings and network-level packet drops, OpenClaw can run arbitrary web crawlers, dynamic shell tasks, and autonomous tool workflows around the clock without compromising the underlying host system.
Deploying OpenClaw with Docker Compose and API Provider Bridges
Moving from container sandboxing to multi-service orchestration requires decoupling the core agent engine from state persistence and LLM API routing. Running openclaw on vps 24/7 requires an architecture resilient to upstream API rate limits, transport timeouts on long-running streaming connections, and silent network disconnects.
The production topology deploys three tightly bound services: 1. OpenClaw Core (openclaw-gateway): The orchestration daemon executing agent loops, maintaining runtime context, and driving the sandboxed runners. 2. Inference Bridge (litellm-proxy): An egress gateway that standardizes requests across OpenAI, Anthropic, and local inference backends (vLLM / Ollama), enforcing retries, backoff, and model fallbacks. 3. State & Queue Cache (valkey-state): High-throughput in-memory storage (Valkey/Redis drop-in) managing task states, tool execution locks, and semantic response caching to prevent duplicate API spend.
┌─────────────────────────────────────────────────────┐
│ tropic.host KVM Node │
│ Linux Kernel 6.8+ (BBR + cgroups v2) │
└──────────────────────────┬──────────────────────────┘
│
┌───────────────────────────────┴───────────────────────────────┐
│ Isolated Docker Bridge Network │
│ │
┌──────────────▼─────────────┐ gRPC / HTTP ┌─────────────────────────────┐ │
│ openclaw-gateway ├─────────────────► litellm-proxy │ │
│ (Core Autonomous Loop) │ │ (Routing, Retry, Fallback)│ │
└──────────────┬─────────────┘ └──────────────┬──────────────┘ │
│ │ │
│ Fast State / Locks │ HTTPS (BBR) │
┌──────────────▼─────────────┐ │ │
│ valkey-state │ │ │
│ (Memory & Task Queues) │ │ │
└────────────────────────────┘ │ │
│ │
└───────────────────────────────┬──────────────┘ │
│ │
▼ │
┌───────────────────────────────────────────────┐ │
│ External API Providers │ │
│ Anthropic API (Claude 3.7 Sonnet) │ │
│ OpenAI API (GPT-4o / o3-mini) │ │
│ Local / Private vLLM (Self-hosted Fallback) │ │
└───────────────────────────────────────────────┘ │
Production Docker Compose Stack
The stack is defined in /opt/openclaw/docker-compose.yml. It applies cgroups v2 resource ceilings, drops non-essential Linux capabilities, mounts configuration files in read-only mode, and establishes log-rotation policies to prevent unbounded disk growth during continuous autonomous runs.
version: "3.8"
networks:
openclaw-mesh:
driver: bridge
ipam:
driver: default
config:
- subnet: 172.28.10.0/24
driver_opts:
com.docker.network.bridge.name: br-openclaw
volumes:
valkey_data:
driver: local
openclaw_workspace:
driver: local
services:
valkey-state:
image: valkey/valkey:7.2-alpine
container_name: openclaw-valkey
restart: unless-stopped
command: >
valkey-server
--save 300 1
--maxmemory 512mb
--maxmemory-policy allkeys-lru
--appendonly yes
--tcp-backlog 511
--protected-mode yes
--requirepass "${VALKEY_PASSWORD}"
networks:
- openclaw-mesh
volumes:
- valkey_data:/data
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
cap_add:
- SETUID
- SETGID
- CHOWN
deploy:
resources:
limits:
cpus: "1.00"
memory: 768M
reservations:
cpus: "0.25"
memory: 256M
logging:
driver: "json-file"
options:
max-size: "20m"
max-file: "5"
healthcheck:
test: ["CMD", "valkey-cli", "-a", "${VALKEY_PASSWORD}", "ping"]
interval: 10s
timeout: 3s
retries: 5
start_period: 5s
litellm-proxy:
image: ghcr.io/berriai/litellm:main-v1.45.0
container_name: openclaw-litellm
restart: unless-stopped
environment:
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
- DATABASE_URL=redis://:${VALKEY_PASSWORD}@valkey-state:6379/0
- STORE_MODEL_IN_DB=False
networks:
- openclaw-mesh
volumes:
- ./config/litellm.yaml:/app/config.yaml:ro
command: ["--config", "/app/config.yaml", "--port", "4000", "--num_workers", "4"]
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
deploy:
resources:
limits:
cpus: "2.00"
memory: 1536M
reservations:
cpus: "0.50"
memory: 512M
logging:
driver: "json-file"
options:
max-size: "50m"
max-file: "5"
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/liveliness')"]
interval: 15s
timeout: 5s
retries: 3
start_period: 15s
openclaw-gateway:
image: openclaw/gateway:v1.2.4
container_name: openclaw-gateway
restart: unless-stopped
depends_on:
valkey-state:
condition: service_healthy
litellm-proxy:
condition: service_healthy
env_file:
- .env
environment:
- OPENCLAW_ENV=production
- REDIS_URL=redis://:${VALKEY_PASSWORD}@valkey-state:6379/1
- LLM_API_BASE=http://litellm-proxy:4000
- LLM_API_KEY=${LITELLM_MASTER_KEY}
- HTTP_CLIENT_TIMEOUT=120
- MAX_CONCURRENT_TASKS=16
networks:
- openclaw-mesh
ports:
- "127.0.0.1:8080:8080"
volumes:
- ./config/openclaw.json:/etc/openclaw/openclaw.json:ro
- openclaw_workspace:/var/lib/openclaw/workspace
security_opt:
- no-new-privileges:true
deploy:
resources:
limits:
cpus: "4.00"
memory: 4096M
reservations:
cpus: "1.00"
memory: 1024M
logging:
driver: "json-file"
options:
max-size: "100m"
max-file: "7"
healthcheck:
test: ["CMD", "curl", "-fsS", "http://127.0.0.1:8080/health"]
interval: 10s
timeout: 4s
retries: 3
start_period: 10s
Configuring the API Provider Bridge (litellm.yaml)
Routing queries through an internal bridge provides three critical advantages for 24/7 autonomous agents: 1. Dynamic Fallbacks: Seamless transition from Anthropic Claude 3.7 Sonnet down to OpenAI GPT-4o or a self-hosted vLLM instance when encountering upstream HTTP 429 (Rate Limit) or HTTP 529 (Overloaded). 2. Unified Semantic Caching: Shared state across tasks using the internal Valkey backend reduces token consumption for repetitive system instructions and codebase analysis. 3. Strict Egress Timeouts and Exponential Backoff: Eliminates hung worker threads when upstream Server-Sent Events (SSE) streams stall mid-generation.
Save the routing rules inside /opt/openclaw/config/litellm.yaml:
model_list:
- model_name: agent-reasoning
litellm_params:
model: anthropic/claude-3-7-sonnet-20250219
api_key: "os.environ/ANTHROPIC_API_KEY"
timeout: 120
max_retries: 3
rpm: 1000
stream: true
- model_name: agent-reasoning-fallback
litellm_params:
model: openai/gpt-4o
api_key: "os.environ/OPENAI_API_KEY"
timeout: 90
max_retries: 3
rpm: 2000
stream: true
- model_name: agent-fast
litellm_params:
model: openai/gpt-4o-mini
api_key: "os.environ/OPENAI_API_KEY"
timeout: 30
max_retries: 2
rpm: 5000
- model_name: agent-local-fallback
litellm_params:
model: openai/local-deepseek-r1
api_base: "http://10.0.0.5:8000/v1"
api_key: "EMPTY"
timeout: 180
max_retries: 1
router_settings:
routing_strategy: "latency-based-routing"
redis_host: "valkey-state"
redis_port: 6379
redis_password: "os.environ/VALKEY_PASSWORD"
fallbacks:
- agent-reasoning: ["agent-reasoning-fallback", "agent-local-fallback"]
allowed_fails: 2
cooldown_time: 30
general_settings:
master_key: "os.environ/LITELLM_MASTER_KEY"
completion_cache: true
cache_type: "redis"
Generate the corresponding /opt/openclaw/.env file:
install -m 600 /dev/null /opt/openclaw/.env
cat <<EOF > /opt/openclaw/.env
VALKEY_PASSWORD=$(openssl rand -hex 24)
LITELLM_MASTER_KEY=sk-proxy-$(openssl rand -hex 24)
ANTHROPIC_API_KEY=sk-ant-api03-...
OPENAI_API_KEY=sk-proj-...
EOF
Kernel and Socket Tuning for p99 Latency and Streaming Resiliency
Autonomous agents executing on an unoptimized Linux host frequently suffer from socket exhaustion (TIME_WAIT buckets filling up), broken pipe exceptions on long SSE streaming tokens, and context switching latency under burst load.
On tropic.host KVM instances, CPU oversubscription is strictly forbidden. With dedicated vCPU scheduling delivering guaranteed %st = 0.0% (CPU Steal Time) on AMD EPYC and Ryzen 9 nodes, CPU scheduler delays will not stall streaming parser loops. However, the Linux network stack must be explicitly configured to prevent connection timeouts when maintaining hundreds of long-lived HTTP/2 streams across transcontinental cloud regions.
Apply the following production parameters in /etc/sysctl.d/99-openclaw-network.conf:
# Enforce TCP BBR for rapid congestion recovery over long-haul transit
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr
# Prevent socket starvation and connection drops under high concurrency
net.core.somaxconn = 8192
net.ipv4.tcp_max_syn_backlog = 8192
net.core.netdev_max_backlog = 16384
# Accelerate reuse of TIME_WAIT sockets for outgoing proxy connections
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
# Mitigate half-dead streaming connection timeouts
# Default Linux 7200s keepalive will leak stale sockets through NAT/firewalls
net.ipv4.tcp_keepalive_time = 60
net.ipv4.tcp_keepalive_intvl = 10
net.ipv4.tcp_keepalive_probes = 5
# Disable slow start restart after idle to maintain high throughput on burst calls
net.ipv4.tcp_slow_start_after_idle = 0
# Increase local port range for high-rate API client polling
net.ipv4.ip_local_port_range = 10240 65535
# File descriptor maximums across the entire OS
fs.file-max = 2097152
Activate the sysctl parameters immediately without a reboot:
sysctl --system
Verify that BBR is actively handling congestion control:
sysctl net.ipv4.tcp_congestion_control
# Expected output: net.ipv4.tcp_congestion_control = bbr
lsmod | grep bbr
# Expected output: tcp_bbr 20480 1
Enterprise PCIe 4.0 NVMe drives on tropic.host deliver random 4K QD1 reads exceeding 50,000 IOPS. This hardware layer ensures that Valkey's write-ahead AOF disk flushes (appendfsync everysec) and Docker container JSON log pipelines commit immediately without introducing I/O wait spikes (%iowait < 0.1%), keeping overall API bridge transaction latency strictly bounded at p99 levels under 120 ms.
Initializing and Validating the Deployment
- Create the operational directory tree and fix filesystem permissions:
mkdir -p /opt/openclaw/{config,workspace}
chmod 700 /opt/openclaw
chmod 600 /opt/openclaw/.env
- Launch the stack in detached mode:
cd /opt/openclaw
docker compose up -d
- Verify container health check statuses:
docker compose ps
Expected output:
NAME IMAGE COMMAND SERVICE CREATED STATUS PORTS
openclaw-gateway openclaw/gateway:v1.2.4 "/entrypoint.sh ..." openclaw-gateway 1 minute ago Up 1 minute (healthy) 127.0.0.1:8080->8080/tcp
openclaw-litellm ghcr.io/berriai/litellm:main-v1.45.0"litellm --config /a…" litellm-proxy 1 minute ago Up 1 minute (healthy) 4000/tcp
openclaw-valkey valkey/valkey:7.2-alpine "docker-entrypoint.s…" valkey-state 1 minute ago Up 1 minute (healthy) 6379/tcp
- Verify downstream streaming and failover routing directly through the bridge:
Test that the bridge responds over HTTP with chunked Server-Sent Events (SSE) without protocol buffering:
source /opt/openclaw/.env
curl -N -X POST http://127.0.0.1:4000/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${LITELLM_MASTER_KEY}" \
-d '{
"model": "agent-reasoning",
"messages": [
{"role": "system", "content": "You are a deterministic system health verifier. Return PONG."},
{"role": "user", "content": "PING"}
],
"stream": true,
"max_tokens": 10
}'
Expected stream chunk output:
data: {"id":"chatcmpl-9x1","object":"chat.completion.chunk","created":1741100000,"model":"claude-3-7-sonnet-20250219","choices":[{"index":0,"delta":{"role":"assistant","content":"P"},"finish_reason":null}]}
data: {"id":"chatcmpl-9x1","object":"chat.completion.chunk","created":1741100000,"model":"claude-3-7-sonnet-20250219","choices":[{"index":0,"delta":{"content":"ONG"},"finish_reason":null}]}
data: {"id":"chatcmpl-9x1","object":"chat.completion.chunk","created":1741100000,"model":"claude-3-7-sonnet-20250219","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
- Test simulated upstream failure and routing fallbacks:
To ensure 24/7 unassisted autonomy, invalidate the Anthropic token temporarily in LiteLLM or inject an upstream network block. Send the same query:
docker exec -it openclaw-litellm litellm --test-fallback
Inspect the proxy logs to verify that the request automatically routes to the secondary provider (openai/gpt-4o) within 250 ms:
docker logs openclaw-litellm --tail 50 | grep -E "(Fallback|Switching|HTTP Error)"
Expected log traces:
2026-10-04 13:52:11,402 - litellm.router - WARNING - Received error from anthropic/claude-3-7-sonnet-20250219: 429 RateLimitError. Attempting fallback 1/2...
2026-10-04 13:52:11,405 - litellm.router - INFO - Routing request to fallback model: openai/gpt-4o
2026-10-04 13:52:11,620 - litellm.proxy - INFO - 200 OK | Model: openai/gpt-4o | Latency: 215ms | Response complete
With the isolated bridge handling connection recycling, token-bucket rate limits, and fallback switching over low-latency BBR sockets, OpenClaw maintains an uninterrupted execution pipeline 24/7 regardless of upstream third-party service degradation.
Continuous Task Scheduling, Headless Browser Drivers, and Webhooks
Operating openclaw on vps 24/7 demands an autonomous execution runtime capable of interacting with external interfaces beyond plain REST APIs. Production agent workflows inevitably encounter dynamic web surfaces (single-page applications requiring JavaScript hydration), scheduled recurring operations (data aggregation, state audits, health telemetry), and external push events (inbound webhook triggers).
Executing headless browsers and persistent workers inside continuous KVM instances introduces non-trivial OS-level failure modes: memory fragmentation from orphan browser renderers, shared-memory buffer depletion, and process starvation. Stabilizing this layer requires explicit Linux kernel parameterization, container-level process isolation, and a decoupled event-driven architecture.
Kernel Tuning and Shared Memory Sizing for Headless Chromium
Headless browser automation via Playwright relies heavily on Chromium’s multi-process architecture. Every target tab initiates separate renderer, GPU, and network processes that communicate via shared memory segments (/dev/shm).
In standard containerized environments, the default Docker /dev/shm size is restricted to 64 MB. Under continuous scraping loads or complex DOM evaluations, Chromium exhausts this allocation within minutes, manifesting as erratic SIGBUS crashes, browser tab target crashes (TargetClosedError), and rendering hangs.
Before launching headless browser workers, calibrate the host's virtual memory limits and process handles in /etc/sysctl.d/99-browser-runtime.conf:
# /etc/sysctl.d/99-browser-runtime.conf
# Increase maximum system-wide process identifier allocation
kernel.pid_max = 4194304
# Accommodate V8 memory allocations and WebAssembly heap spaces
vm.max_map_count = 262144
# Expand system-wide file descriptor limit to prevent socket/file exhaustion under parallel loads
fs.file-max = 2097152
# Enforce immediate memory reclamation over swap thrashing
vm.swappiness = 10
vm.overcommit_memory = 1
Apply the parameters immediately:
sudo sysctl --system
To prevent runaway Chromium processes from spawning fork bombs or consuming unbounded system threads, configure explicit cgroups v2 process count limits and expand /dev/shm directly within your container definitions using Docker Compose.
Production Containerization: Playwright Worker with Init Zombie Reaping
When Chromium processes terminate abnormally inside containers, the Linux kernel re-parents orphan child threads to PID 1. If PID 1 is a standard Node.js or Python process rather than a designated init system, it fails to execute the requisite waitpid() syscalls. Over days of continuous operation, the process table fills with <defunct> zombie processes, ultimately breaking system-level fork operations.
We mitigate this by deploying a dedicated Playwright automation service utilizing a dumb-init/Tini reaper (init: true), explicit memory bounds, and a loopback-isolated WebSocket communication channel for the OpenClaw orchestrator.
Here is the production-grade worker topology in docker-compose.automation.yml:
version: "3.8"
services:
openclaw-playwright:
image: mcr.microsoft.com/playwright:v1.49.0-jammy
container_name: openclaw-playwright
restart: unless-stopped
init: true # Crucial: invokes Tini to reap zombie Chromium renderer processes
ipc: host # Alternatively allocate dedicated tmpfs via shm_size below
shm_size: "2gb" # Allocates sufficient memory for heavy DOM trees and canvas rendering
security_opt:
- seccomp=unconfined # Required for Chrome sandbox operations without granting root
environment:
- PLAYWRIGHT_BROWSERS_PATH=/ms-playwright
- PORT=3000
pids_limit: 2048 # cgroups v2: restricts max thread/process exhaustion
deploy:
resources:
limits:
cpus: "2.0"
memory: 4096M
reservations:
cpus: "1.0"
memory: 2048M
networks:
openclaw-internal:
ipv4_address: 172.28.0.25
command: ["sh", "-c", "npx playwright run-server --port 3000 --path /browser"]
networks:
openclaw-internal:
external: true
The OpenClaw core engine connects to this isolated browser pool over the local software bridge network (ws://172.28.0.25:3000/browser).
When provisioning your runtime node on a high-frequency tropic.host KVM instance, the presence of 0.0% CPU Steal Time (%st = 0.0%) and enterprise NVMe arrays (random 4K QD1 reads exceeding 50,000 IOPS) becomes critical here. Modern dynamic web applications execute heavy client-side JavaScript bundling. On oversubscribed virtual environments, vCPU throttling during CSS/DOM parsing spikes navigation p99 latency beyond 8,000 ms, triggering unrecoverable Playwright connection timeouts. Dedicated AMD EPYC and Ryzen 9 cores guarantee deterministic DOM evaluation with consistent sub-350 ms render loops.
Robust Playwright Session Management via Python
To execute tasks reliably over weeks of unattended uptime, client scripts must implement deterministic session tearing, context disposal, and stealth viewport parameters.
Below is the production browser driver module (browser_driver.py) utilized by OpenClaw:
"""
OpenClaw Automated Browser Engine - 24/7 Context Isolation Driver
"""
import asyncio
import logging
from contextlib import asynccontextmanager
from playwright.async_api import async_playwright, Browser, BrowserContext, Page
logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s")
logger = logging.getLogger("openclaw.browser")
PLAYWRIGHT_WS_ENDPOINT = "ws://172.28.0.25:3000/browser"
class BrowserDriverPool:
def __init__(self, endpoint: str = PLAYWRIGHT_WS_ENDPOINT):
self.endpoint = endpoint
self._playwright = None
self._browser: Browser = None
async def initialize(self):
self._playwright = await async_playwright().start()
# Connect over internal Docker network to the sandboxed Playwright daemon
self._browser = await self._playwright.chromium.connect(self.endpoint)
logger.info("Connected to headless Playwright worker cluster.")
@asynccontextmanager
async def session(self) -> Page:
"""
Yields an isolated, transient browser context to prevent cross-run cookie/cache pollution.
Guarantees forceful context destruction on task completion or unhandled exceptions.
"""
if not self._browser or not self._browser.is_connected():
await self.initialize()
context: BrowserContext = await self._browser.new_context(
viewport={"width": 1920, "height": 1080},
user_agent="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36",
locale="en-US",
timezone_id="UTC",
ignore_https_errors=False
)
# Enforce hard request timeouts to prevent hung socket connections from blocking worker pools
context.set_default_timeout(20000)
context.set_default_navigation_timeout(25000)
page: Page = await context.new_page()
try:
yield page
finally:
await page.close()
await context.close()
async def shutdown(self):
if self._browser:
await self._browser.close()
if self._playwright:
await self._playwright.stop()
logger.info("Playwright automation pool gracefully terminated.")
Deterministic Background Task Scheduling (Systemd Timers vs. Overlapping Crons)
Running background agent routines via standard crontab configurations exposes production systems to race conditions: if a scraping routine or long-running inference job takes 7 minutes due to external rate limiting, a 5-minute cron triggers a secondary overlapping instance. This cascading concurrency quickly leads to memory exhaustion and API ban triggers.
To establish resilient 24/7 background execution on your Linux VPS host, employ systemd timer units paired with non-blocking kernel file locking (flock).
1. The Runner Wrapper Script with Exclusive Locking
Create /usr/local/bin/openclaw-runner.sh:
#!/usr/bin/env bash
set -euo pipefail
LOCK_FILE="/var/run/openclaw_worker.lock"
EXEC_BIN="/opt/openclaw/venv/bin/python"
TARGET_SCRIPT="/opt/openclaw/app/worker.py"
# Obtain non-blocking exclusive file lock (FD 200)
exec 200>"$LOCK_FILE"
if ! flock -n 200; then
echo "[$(date -u +'%Y-%m-%dT%H:%M:%SZ')] Worker execution skipped: Prior task instance still active." >&2
exit 0
fi
# Execute target workload within the isolated environment
"$EXEC_BIN" "$TARGET_SCRIPT" --mode continuous --max-iterations 50
Ensure correct permissions:
sudo chmod +x /usr/local/bin/openclaw-runner.sh
2. Systemd Service Definition
Create /etc/systemd/system/openclaw-scheduler.service:
[Unit]
Description=OpenClaw Autonomous Background Worker Execution
After=network.target docker.service
Wants=openclaw-playwright.service
[Service]
Type=oneshot
User=openclaw
Group=openclaw
WorkingDirectory=/opt/openclaw
ExecStart=/usr/local/bin/openclaw-runner.sh
StandardOutput=append:/var/log/openclaw/scheduler.log
StandardError=append:/var/log/openclaw/scheduler-err.log
# Enforce execution timeout; terminate tasks exceeding 15 minutes
TimeoutStartSec=900
# Security sandboxing
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/var/log/openclaw /var/run /opt/openclaw/data
NoNewPrivileges=true
PrivateTmp=true
[Install]
WantedBy=multi-user.target
3. Systemd High-Precision Monotonic Timer
Create /etc/systemd/system/openclaw-scheduler.timer:
[Unit]
Description=OpenClaw 10-Minute Monotonic Execution Timer
Requires=openclaw-scheduler.service
[Timer]
# Fire 2 minutes after boot, then precisely every 10 minutes relative to the last activation
OnBootSec=2min
OnUnitActiveSec=10min
AccuracySec=100ms
Persistent=true
[Install]
WantedBy=timers.target
Reload and activate:
sudo systemctl daemon-reload
sudo systemctl enable --now openclaw-scheduler.timer
Verify timer precision and next scheduled trigger:
systemctl list-timers --all | grep openclaw
Inbound Webhook Ingress and Cryptographic Verification
Periodic polling can be inefficient for external events (e.g., alert dispatchers, code repository updates, payment notifications). OpenClaw integrates an asynchronous webhook listener built on FastAPI and Uvicorn, secured behind an HMAC-SHA256 signature verification pipeline.
The webhook receiver executes on the internal Docker network, exposed externally via a TLS-terminating reverse proxy (Nginx or Caddy) over high-throughput, BBR-optimized TCP sockets.
Here is the production implementation of /opt/openclaw/app/webhook_listener.py:
"""
OpenClaw Production Webhook Ingress Receiver
Implements constant-time HMAC-SHA256 payload verification.
"""
import hmac
import hashlib
import os
from fastapi import FastAPI, Request, Header, HTTPException, status, BackgroundTasks
import uvicorn
WEBHOOK_SECRET = os.environ.get("OPENCLAW_WEBHOOK_SECRET", "").encode("utf-8")
if not WEBHOOK_SECRET:
raise RuntimeError("FATAL: OPENCLAW_WEBHOOK_SECRET environment variable is unset.")
app = FastAPI(title="OpenClaw Ingress", docs_url=None, redoc_url=None)
def verify_hmac_signature(payload: bytes, signature_header: str) -> bool:
"""
Validates HMAC signature using constant-time comparison to prevent timing attack vectors.
Expected signature header format: 'sha256=<hex_digest>'
"""
if not signature_header or not signature_header.startswith("sha256="):
return False
provided_signature = signature_header.split("sha256=")[-1].strip()
computed_signature = hmac.new(WEBHOOK_SECRET, payload, hashlib.sha256).hexdigest()
return hmac.compare_digest(provided_signature, computed_signature)
async def dispatch_autonomous_pipeline(payload: dict):
"""
Asynchronous target task invoked via FastAPI BackgroundTasks.
Frees the HTTP connection immediately while delegating processing to the agent loop.
"""
# Pipeline invocation logic (e.g., dispatching into Redis queue or Celery worker)
pass
@app.post("/api/v1/webhook", status_code=status.HTTP_202_ACCEPTED)
async def handle_incoming_webhook(
request: Request,
background_tasks: BackgroundTasks,
x_hub_signature_256: str = Header(None, alias="X-Hub-Signature-256")
):
raw_body = await request.body()
if not verify_hmac_signature(raw_body, x_hub_signature_256):
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail="Invalid cryptographic signature."
)
try:
json_data = await request.json()
except Exception:
raise HTTPException(
status_code=status.HTTP_400_BAD_REQUEST,
detail="Malformed JSON body."
)
# Hand off payload to background execution pool to maintain <15ms HTTP 202 acknowledgment
background_tasks.add_task(dispatch_autonomous_pipeline, json_data)
return {"status": "accepted", "queued": True}
if __name__ == "__main__":
uvicorn.run(app, host="127.0.0.1", port=8080, log_level="warning", access_log=False)
Pair this with an Nginx host ingress block terminating TLS 1.3 with upstream keepalive buffers:
# /etc/nginx/conf.d/openclaw_ingress.conf
upstream openclaw_webhook {
server 127.0.0.1:8080;
keepalive 32;
}
server {
listen 443 ssl http2;
server_name hook.yourdomain.internal;
ssl_certificate /etc/letsencrypt/live/hook.yourdomain.internal/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/hook.yourdomain.internal/privkey.pem;
ssl_protocols TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
client_max_body_size 5M;
client_body_buffer_size 128k;
location /api/v1/webhook {
proxy_pass http://openclaw_webhook;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_connect_timeout 5s;
proxy_read_timeout 10s;
proxy_send_timeout 10s;
}
}
Through hardened Playwright driver sandboxing, non-overlapping monotonic timer synchronization, and constant-time webhook ingress validation, your OpenClaw agent deployment operates as a durable, self-healing background automation engine across multi-week continuous cycles.
API Key Security, Rate Limiting, and Cost Containment Strategies
Unattended, 24/7 execution of OpenClaw exposes two existential operational hazards: credential compromise via browser runtime exploitation and runaway financial exhaustion through recursive prompt-loop failures. When an autonomous agent processes untrusted web pages via headless Chromium, an arbitrary code execution vulnerability or prompt injection payload can turn the crawler into an outbound exfiltration vector. Concurrently, an unthrottled loop consuming OpenAI, Anthropic, or DeepSeek inference endpoints can exhaust monthly billing allowances in minutes.
Mitigating these threat vectors requires zero-trust architecture directly at the host layer: encrypting secrets at rest, isolating memory, implementing egress network filtering via nftables, and terminating outbound inference calls through a local rate-limiting reverse proxy with automated circuit breakers.
Zero-Plaintext Secret Provisioning: SOPS, age, and Ephemeral In-Memory Injection
Storing plaintext API keys inside standard .env files or hardcoded configuration blocks leaves keys vulnerable to host-level process snooping, accidental repository commits, and memory dump harvesting. On Linux hosts, secrets must remain encrypted on disk and be decrypted exclusively into an ephemeral, non-swappable in-memory mount during daemon initialization.
1. Asymmetric Secret Encryption with age and Mozilla SOPS
Generate an isolated, non-interactive age key pair directly on the deployment host:
# Generate the age identity file in an isolated directory
install -d -m 0700 /etc/openclaw/keys
age-keygen -o /etc/openclaw/keys/openclaw.key
chmod 0400 /etc/openclaw/keys/openclaw.key
# Extract the corresponding public key
AGE_PUBKEY=$(age-keygen -y /etc/openclaw/keys/openclaw.key)
Construct your encrypted environment payload (secrets.enc.env) using Mozilla SOPS, binding it to the public key:
sops --encrypt --age "${AGE_PUBKEY}" \
--output /etc/openclaw/secrets.enc.env /dev/stdin <<EOF
OPENAI_API_KEY="sk-proj-prod-live-example994820148102"
ANTHROPIC_API_KEY="sk-ant-api03-live-example8810294"
DEEPSEEK_API_KEY="sk-dsk-live-example102948"
SERP_API_TOKEN="token_prod_serp_991823"
OPENCLAW_HMAC_SECRET="c4ca4238a0b923820dcc509a6f75849b"
EOF
chmod 0400 /etc/openclaw/secrets.enc.env
2. Systemd Dynamic Injection via tmpfs and LoadCredentialEncrypted
Avoid passing sensitive variables directly through docker-compose.yml or global shell profiles. Instead, decrypt secrets dynamically into a dedicated RAM-backed tmpfs volume with read permissions limited strictly to the unprivileged OpenClaw execution user (UID 10001).
# /etc/systemd/system/openclaw-worker.service
[Unit]
Description=OpenClaw Autonomous Core Automation Worker
After=network-online.target local-fs.target
Wants=network-online.target
[Service]
Type=simple
User=openclaw
Group=openclaw
WorkingDirectory=/opt/openclaw
# Mount an ephemeral, non-swappable tmpfs for decrypted credentials
RuntimeDirectory=openclaw/credentials
RuntimeDirectoryMode=0700
# Prevent core dumps and trace snooping
LimitCORE=0
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
PrivateTmp=true
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
RestrictRealtime=true
MemoryDenyWriteExecute=true
# Decrypt directly to ephemeral memory right before container or process spawn
ExecStartPre=/usr/bin/install -d -m 0700 -o openclaw -g openclaw /run/openclaw/credentials
ExecStartPre=/bin/bash -c '/usr/bin/sops --decrypt --age-key-file /etc/openclaw/keys/openclaw.key /etc/openclaw/secrets.enc.env > /run/openclaw/credentials/runtime.env && chmod 0400 /run/openclaw/credentials/runtime.env'
# Execute the supervisor targeting the RAM-resident env
ExecStart=/opt/openclaw/venv/bin/python3 -m openclaw.core --env-file /run/openclaw/credentials/runtime.env
# Clean up memory buffers immediately on termination
ExecStopPost=/bin/rm -f /run/openclaw/credentials/runtime.env
Restart=always
RestartSec=5s
[Install]
WantedBy=multi-user.target
3. Linux Kernel Process Isolation (Hardening /proc)
In standard Linux installations, any local user or compromised low-privilege daemon can inspect environment variables of other running processes via /proc/$PID/environ. Close this attack vector by remounting /proc with hidepid=invisible (hidepid=2) and applying strict Yama ptrace restrictions:
# /etc/fstab entry to restrict /proc visibility
proc /proc proc defaults,nosuid,nodev,noexec,relatime,hidepid=2,gid=openclaw 0 0
Apply immediately and enforce via /etc/sysctl.d/99-security-hardening.conf:
# Prevent arbitrary ptrace attaching across non-child process boundaries
kernel.yama.ptrace_scope = 2
# Disable suid core dumps to eliminate credential extraction from core files
fs.suid_dumpable = 0
# Enforce strict BPF JIT compiler hardening
net.core.bpf_jit_harden = 2
# Restrict dmesg inspection to root
kernel.dmesg_restrict = 1
Run sysctl --system and mount -o remount /proc to enforce the parameters without dropping active user connections.
Strict Egress Filtering via nftables
By default, Docker and container engines configure dynamic iptables chains that allow unconstrained outbound traffic on all ports. If OpenClaw's browser driver accesses an adversarial URL containing prompt injection or remote code execution exploits, the compromised process can immediately pivot, establishing reverse shells or exfiltrating data to arbitrary listener IPs.
Enforce an egress-default-drop policy on the host firewall. Only authenticated TLS sessions (tcp/443) to explicit upstream model endpoints and trusted DNS resolution (udp/53, tcp/53) are permitted.
#!/usr/sbin/nft -f
# /etc/nftables.conf
flush ruleset
table inet filter {
# Define trusted DNS servers (e.g., Cloudflare and Quad9)
set trusted_dns {
type ipv4_addr
elements = { 1.1.1.1, 1.0.0.1, 9.9.9.9 }
}
chain input {
type filter hook input priority filter; policy drop;
# Connection tracking: keep established sessions alive
ct state established,related accept
ct state invalid drop
# Loopback interface
iif "lo" accept
# Host SSH (Restrict to your management subnet in production)
tcp dport 22 ct state new accept
# Ingress Webhook endpoint (Nginx reverse proxy)
tcp dport { 80, 443 } ct state new accept
# ICMP echo-request (rate-limited to 10/sec to prevent ping floods)
icmp type echo-request limit rate 10/second accept
}
chain forward {
type filter hook forward priority filter; policy drop;
# Docker internal bridges forward rules
ct state established,related accept
ct state invalid drop
}
chain output {
type filter hook output priority filter; policy drop;
# Allow local inter-process communication
oif "lo" accept
# Maintain established external connections
ct state established,related accept
ct state invalid drop
# Allow DNS resolution exclusively to explicitly verified upstream resolvers
udp dport 53 ip daddr @trusted_dns accept
tcp dport 53 ip daddr @trusted_dns accept
# Allow NTP time synchronization
udp dport 123 accept
# Allow outbound HTTPS (Port 443) exclusively for API ingestion and LLM calls
tcp dport 443 accept
# Reject all unauthorized egress with TCP reset to avoid process hanging
tcp flags syn / fin,syn,rst,ack reject with tcp reset
reject with icmpx type admin-prohibited
}
}
Load and persist the configuration across reboots:
nft -f /etc/nftables.conf
systemctl enable --now nftables
Local Reverse Proxy Rate Limiting & Token Budget Circuit Breaking
Autonomous loops running 24/7 can cascade into rapid-fire retry storms if upstream APIs return transient HTTP 5xx codes or if an LLM generates recursive tool calls. Direct interaction between OpenClaw and third-party APIs must be strictly prohibited. All outbound model traffic must pass through a local gateway (such as LiteLLM or an Envoy-based proxy) executing on 127.0.0.1, enforcing:
- Per-minute token consumption ceilings (TPM).
- Per-day financial hard caps (USD budget enforcement).
- Sliding-window exponential backoff that halts queries when p99 latency exceeds defined operational thresholds.
┌─────────────────────────┐ Unix Socket / HTTP ┌─────────────────────────┐
│ OpenClaw Agent Core │ ──────────────────────────────> │ LiteLLM Token Gateway │
│ (Headless Orchestrator)│ │ (127.0.0.1:4000) │
└─────────────────────────┘ └────────────┬────────────┘
│
┌─────────────────────────────────┴──────────────────┐
│ Real-Time Spend Tracking & Hard Circuit Breaker │
│ (Local In-Memory Redis Engine: /run/redis.sock) │
└─────────────────────────────────┬──────────────────┘
│
Outbound TLS 1.3 │ Verified API Key
(TCP BBR Accelerated) │ Ingestion Only
▼
┌─────────────────────────┐
│ Upstream Provider API │
│ (OpenAI / Anthropic) │
└─────────────────────────┘
Production LiteLLM Gateway Configuration
Deploy an isolated proxy configuration file:
# /etc/openclaw/proxy_config.yaml
model_list:
- model_name: gpt-4o-pipeline
litellm_params:
model: openai/gpt-4o
api_key: "os.environ/OPENAI_API_KEY"
timeout: 30
max_retries: 2
rpm: 120
tpm: 100000
- model_name: claude-sonnet-pipeline
litellm_params:
model: anthropic/claude-3-5-sonnet-20241022
api_key: "os.environ/ANTHROPIC_API_KEY"
timeout: 35
max_retries: 2
rpm: 80
tpm: 80000
router_settings:
routing_strategy: "latency-based-routing"
redis_host: "127.0.0.1"
redis_port: 6379
enable_pre_call_checks: true
general_settings:
master_key: "sk-local-openclaw-loop-guard-auth-token"
max_budget: 15.00 # Hard daily spend limit: $15.00 USD
budget_duration: "1d" # Budget reset interval
alerting: ["webhook"]
alerting_threshold: 0.90 # Trigger warnings at 90% threshold
litellm_settings:
drop_params: true
set_verbose: false
Redis-Backed Sliding-Window Circuit Breaker Integration
To track rate limits and budgets across multi-threaded workers with deterministic, sub-millisecond overhead, route cache storage through a local Redis instance backed by a Unix domain socket:
# /opt/openclaw/openclaw/security/circuit_breaker.py
import sys
import redis
import httpx
import logging
from typing import Dict, Any
logger = logging.getLogger("openclaw.breaker")
class TokenBudgetExceededException(Exception):
"""Raised when the financial ceiling for the current sliding window is reached."""
pass
class LLMProxyClient:
def __init__(self, proxy_base_url: str, master_token: str, redis_sock: str = "/var/run/redis/redis-server.sock"):
self.client = httpx.Client(
base_url=proxy_base_url,
headers={
"Authorization": f"Bearer {master_token}",
"Content-Type": "application/json"
},
timeout=httpx.Timeout(40.0, connect=5.0)
)
self.r = redis.Redis(unix_socket_path=redis_sock, decode_responses=True)
def dispatch_inference(self, model_alias: str, messages: list, estimated_tokens: int) -> Dict[str, Any]:
# Atomic query to check if circuit breaker has tripped
circuit_status = self.r.get("openclaw:circuit_tripped")
if circuit_status == "1":
logger.critical("LLM Circuit breaker is OPEN. Halting all outbound inference queries.")
raise TokenBudgetExceededException("Daily spend cap reached. Execution suspended by host proxy.")
payload = {
"model": model_alias,
"messages": messages,
"max_tokens": 4096,
"temperature": 0.2
}
try:
response = self.client.post("/v1/chat/completions", json=payload)
# Trap HTTP 429 (Provider Rate Limit) or HTTP 400 with budget error
if response.status_code == 429 or "budget_exceeded" in response.text:
self.trip_circuit_breaker(reason="Provider 429 / Budget Limit Encountered")
raise TokenBudgetExceededException("Upstream budget exceeded or rate limited.")
response.raise_for_status()
data = response.json()
# Record token usage back into Redis sliding-window metrics
usage = data.get("usage", {})
total_tokens = usage.get("total_tokens", estimated_tokens)
self._update_sliding_window_usage(total_tokens)
return data
except httpx.RequestError as exc:
logger.error(f"Inference gateway communication failure: {exc}")
raise
def trip_circuit_breaker(self, reason: str):
# Trip the kill-switch with an active TTL of 24 hours (86400 seconds)
self.r.setex("openclaw:circuit_tripped", 86400, "1")
logger.emergency(f"EMERGENCY CIRCUIT TRIP: {reason}. Manual administrative intervention required.")
# Optional: Dispatch alert via host notification script
sys.exit(88)
def _update_sliding_window_usage(self, token_count: int):
pipe = self.r.pipeline()
pipe.incrby("openclaw:metrics:tokens:hourly", token_count)
pipe.expire("openclaw:metrics:tokens:hourly", 3600)
pipe.execute()
Low-Jitter Infrastructure Foundations on tropic.host
Enforcing cryptographic boundary checks, local proxy parsing, and deterministic token rate-limiting introduces multiple micro-delays into the processing chain. On generic multi-tenant hosting with shared CPU cores, noisy-neighbor contention drives CPU Steal Time (%st) above 5–15%, degrading p99 TLS handshake times with upstream APIs from 45ms to upwards of 450ms. When connections stall, workers queue up, compounding retry overhead and inflating token consumption via redundant attempts.
Host Virtualization Comparison Under 120 RPM Cryptographic Token Proxying:
┌──────────────────────────────┬───────────────────┬───────────────────┬──────────────────────┐
│ Virtualization Platform │ CPU Steal (%st) │ Proxy p99 Latency │ TLS Session Jitter │
├──────────────────────────────┼───────────────────┼───────────────────┼──────────────────────┤
│ Oversubscribed Cloud VPS │ 4.2% – 18.0% │ 320ms – 680ms │ High (±180ms) │
│ tropic.host Dedicated KVM │ 0.0% (Guaranteed) │ 18ms – 32ms │ Negligible (±1.5ms) │
└──────────────────────────────┴───────────────────┴───────────────────┴──────────────────────┘
Deploying OpenClaw on high-frequency KVM instances from tropic.host establishes the compute isolation necessary to run this zero-trust stack 24/7 without performance degradation:
- Dedicated Compute Slices with
%st = 0.0%: Enterprise AMD EPYC and high-clock Ryzen 9 cores guarantee that CPU-intensive encryption tasks (SOPS/agepayload decryption, ChaCha20-Poly1305 stream parsing, and constant-time HMAC validations) execute without CPU starvation or thread stalling. - PCIe 4.0 Enterprise NVMe Storage: Random 4K QD1 metrics exceeding 50,000 IOPS ensure that real-time token tracking via on-disk Redis persistence, Playwright local storage mounts, and temporary
tmpfsbuffer flushes write immediately, avoiding kerneliowaitspikes. - TCP BBR on Clean Dedicated IPv4: Every server instance provisioned on tropic.host features direct BGP routing at central European and global transit exchanges (Frankfurt, Amsterdam, Istanbul). Outbound TLS 1.3 handshakes to model APIs benefit from default-enabled TCP BBR congestion control, shrinking socket setup times and preventing packet loss during continuous 24/7 autonomous ingestion. Clean, unshared IP ranges eliminate unexpected CAPTCHAs and origin-level WAF blocks that disrupt automated workflows.
Systemd Process Watchdogs, Monitoring, and Recovery from Runaway Loops
Autonomous agents operating without human intervention inevitably encounter degenerative run states: socket hangs during upstream model inference, zombie browser instances spawning from headless Playwright sessions, unhandled promise rejections, and infinite tool-execution loops. Maintaining continuous reliability for openclaw on vps 24/7 requires shifting process governance from fragile userland process supervisors like PM2 or screen/tmux wrappers to native Linux kernel subsystems.
By anchoring OpenClaw to systemd backed by Linux cgroups v2, you establish deterministic supervision: hardware/software watchdog integration, hard compute and memory boundaries, autonomous crash dumping, and high-throughput journal auditing.
┌────────────────────────────────────────────────────────────────────────┐
│ Linux Kernel / cgroups v2 │
│ [ CPUQuota=200% ] [ MemoryHigh=3.5G ] [ TasksMax=256 (Anti-Fork) ]│
└───────────────────────────────────┬────────────────────────────────────┘
│ Enforces hard resource ceilings
▼
┌────────────────────────────────────────────────────────────────────────┐
│ systemd (PID 1 Supervisor) │
│ Type=notify ───► Unix Socket: /run/systemd/notify │
│ WatchdogSec=30s ◄── (sd_notify "WATCHDOG=1" every 10s) │
└──────────────┬──────────────────────────────────────────▲──────────────┘
│ │ Heartbeat Ping
│ SIGABRT / SIGKILL on timeout │ (Zero-st on
▼ │ tropic.host)
┌─────────────────────────────────────────────────────────┴──────────────┐
│ OpenClaw Runtime Daemon (Worker) │
│ Tool Loop Iteration Counter ──► Circuit Breaker (Max 15 ops/min) │
│ Headless Chromium Context ──► Isolated PID subtree in sandbox │
└────────────────────────────────────────────────────────────────────────┘
1. Hardened Systemd Service Unit Specification
Standard systemd services using Type=simple or Type=forking are blind to deadlocks; a process spinning inside an infinite LLM query loop or frozen on a dropped TCP socket remains active (running) from systemd’s perspective indefinitely.
To achieve self-healing, OpenClaw must be configured as Type=notify with an active software watchdog (WatchdogSec=30s). The daemon signals its health by sending heartbeat pings (WATCHDOG=1) across the $NOTIFY_SOCKET Unix domain socket. If the event loop stalls or deadlocks and fails to notify systemd within the 30-second window, PID 1 issues a SIGABRT to produce a core dump, cleans up the cgroup, and restarts the service.
Create the hardened unit file at /etc/systemd/system/openclaw.service:
[Unit]
Description=OpenClaw Autonomous Agent Daemon
After=network-online.target local-fs.target
Wants=network-online.target
Documentation=https://github.com/openclaw/openclaw
[Service]
Type=notify
User=openclaw
Group=openclaw
WorkingDirectory=/opt/openclaw
EnvironmentFile=/etc/openclaw/openclaw.env
ExecStart=/opt/openclaw/bin/openclaw-runner --config /etc/openclaw/config.yaml
ExecReload=/bin/kill -HUP $MAINPID
# Watchdog and Restart Policies
WatchdogSec=30s
Restart=always
RestartSec=5s
StartLimitIntervalSec=300s
StartLimitBurst=5
TimeoutStartSec=60s
TimeoutStopSec=30s
# Sandboxing and Security Lockdown
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/opt/openclaw/data /opt/openclaw/logs /tmp
PrivateTmp=true
PrivateDevices=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=true
RestrictSUIDSGID=true
RestrictRealtime=true
RestrictNamespaces=true
LockPersonality=true
# cgroups v2 Resource Governance
CPUAccounting=true
CPUQuota=200%
CPUWeight=100
MemoryAccounting=true
MemoryHigh=3.5G
MemoryMax=4.0G
MemorySwapMax=0
TasksAccounting=true
TasksMax=256
DevicePolicy=closed
# Crash Diagnostics
ExecStopPost=/usr/local/bin/openclaw-crash-reporter.sh %p %i %s $EXIT_STATUS
[Install]
WantedBy=multi-user.target
Deterministic Scheduling and Zero Steal Time
Software watchdogs that sample state every 10–15 seconds require deterministic CPU scheduling. On oversubscribed virtualization platforms where physical CPUs are multiplexed among competing tenants, thread scheduling stalls frequently induce CPU Steal Time (%st > 5.0%). When the hypervisor de-schedules the virtual core running OpenClaw's event loop, the process misses its WATCHDOG=1 submission window. The operating system misinterprets this compute starvation as an internal software deadlock, triggering continuous SIGABRT restart storms.
Running the stack on tropic.host prevents this architectural failure mode. Clean KVM hypervisor slices allocate dedicated AMD EPYC and Ryzen 9 vCPUs with guaranteed %st = 0.0%. Kernel interrupts and systemd watchdog verification execute within microsecond tolerances, ensuring that restarts occur strictly upon legitimate runtime deadlocks rather than platform-induced scheduler jitter.
2. Runtime Watchdog Heartbeat and Runaway Loop Circuit Breakers
A watchdog is only as effective as the integrity of the heartbeat logic. If the heartbeat is placed on an uncoupled background thread, it will continue pulsing WATCHDOG=1 even if the agent's primary decision loop is blocked or spinning in an infinite recursive tool call.
The agent's internal engine must link the notification call directly to the progression of its core asynchronous state machine. When deploying Python or Node-based implementations of OpenClaw, integrate the notification directly via systemd-python or through an IPC wrapper interacting with /run/systemd/notify.
Application-Level Watchdog Implementation
The runtime must implement two safeguards: 1. Liveness Heartbeat: Sends WATCHDOG=1 only when the primary loop completes a tick and external health probes (local storage access, Redis session queue) pass. 2. Loop Iteration Circuit Breaker: Detects recursive hallucinations (e.g., an agent executing identical filesystem lookups or CLI queries $> 15$ times within a 60-second window) and intentionally trips the watchdog by halting notifications and writing an emergency diagnostic trace to disk.
#!/usr/bin/env python3
"""
OpenClaw Watchdog & Circuit-Breaker Integration Engine
Location: /opt/openclaw/lib/watchdog_integration.py
"""
import os
import sys
import time
import socket
import logging
from collections import deque
logger = logging.getLogger("openclaw.watchdog")
class SystemdWatchdogClient:
def __init__(self, max_actions_per_window=15, window_seconds=60):
self.notify_socket = os.getenv("NOTIFY_SOCKET")
self.max_actions = max_actions_per_window
self.window_seconds = window_seconds
self.action_history = deque()
self.last_ping = 0.0
if not self.notify_socket:
logger.warning("NOTIFY_SOCKET not detected. Watchdog running in bypass mode.")
def _send(self, payload: bytes):
if not self.notify_socket:
return
# Systemd notification abstract or filesystem Unix domain socket
sock_addr = self.notify_socket
if sock_addr.startswith("@"):
sock_addr = "\0" + sock_addr[1:]
with socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM) as sock:
try:
sock.connect(sock_addr)
sock.sendall(payload)
except OSError as err:
logger.error(f"Failed to transmit sd_notify payload: {err}")
def notify_ready(self):
"""Signals systemd that initialization is complete."""
self._send(b"READY=1\nSTATUS=Worker operational. Initializing agent loop.")
def check_circuit_breaker(self, tool_signature: str) -> bool:
"""
Tracks repetitive tool invocations to identify runaway LLM execution loops.
Returns False if the loop exceeds safety thresholds.
"""
now = time.monotonic()
self.action_history.append((now, tool_signature))
# Evict actions outside sliding window
while self.action_history and self.action_history[0][0] < (now - self.window_seconds):
self.action_history.popleft()
# Count identical tool execution signatures
recent_matches = sum(1 for _, sig in self.action_history if sig == tool_signature)
if recent_matches > self.max_actions:
logger.critical(
f"CIRCUIT BREAKER TRIPPED: Tool '{tool_signature}' executed {recent_matches} "
f"times in {self.window_seconds}s. Starving watchdog to force supervisor restart."
)
# Send status update before silence
self._send(f"STATUS=CRITICAL: Loop detected on {tool_signature}. Aborting.".encode())
return False
return True
def ping_watchdog(self):
"""Transmits WATCHDOG=1 heartbeat if interval requirement is met."""
now = time.monotonic()
# Enforce rate-limit: ping every 10 seconds (for a 30-second WatchdogSec)
if now - self.last_ping >= 10.0:
self._send(b"WATCHDOG=1")
self.last_ping = now
3. cgroups v2 Kernel Resource Fencing
Unchecked headless browser instances (Chromium, Firefox) orchestrated by OpenClaw can experience DOM reference leaks, leading to progressive RAM exhaustion that threatens the host OS. Rather than relying on the Linux generic oom-killer to randomly terminate processes based on badness heuristics, configure strict cgroups v2 memory throttling directly within the unit file:
MemoryHigh=3.5G: Acts as the soft reclaim ceiling. When OpenClaw's memory footprint breaches 3.5 GB, the Linux kernel aggressively reclaims page cache and throttles allocation calls from the agent's threads. This creates a backpressure warning without killing active jobs.MemoryMax=4.0G: The absolute hard ceiling. Exceeding this boundary immediately triggers an out-of-memory termination localized strictly within/system.slice/openclaw.service, leaving SSH daemons, telemetry exporters, and the root system unaffected.MemorySwapMax=0: Completely eliminates swap traversal for the cgroup. Swapping dynamic agent heaps onto disk induces massive page fault thrashing, driving IO wait (%wa) up and stalling the event loop past the 30-second watchdog limit.TasksMax=256: Hard ceiling on total process IDs (threads + child processes). If an error loop cascades and continuously spawns zombie headless browser engines without garbage collection, the kernel prevents process table exhaustion by returning-EAGAINon subsequentfork()orclone()system calls.
4. Post-Mortem Crash Reporting and Diagnostic Script
When systemd triggers a watchdog restart or an OOM kill occurs, forensic context must be captured instantly before the cgroup memory space is reinitialized. The ExecStopPost directive executes independently of whether OpenClaw exits cleanly, aborts, or is terminated via signals.
Deploy /usr/local/bin/openclaw-crash-reporter.sh:
#!/usr/bin/env bash
set -euo pipefail
SERVICE_NAME="${1:-openclaw}"
INSTANCE_NAME="${2:-}"
EXIT_STATUS="${3:-unknown}"
EXIT_CODE="${4:-0}"
TIMESTAMP=$(date -u +"%Y%m%d_%H%M%SZ")
DUMP_DIR="/opt/openclaw/logs/crashes"
DUMP_FILE="${DUMP_DIR}/${SERVICE_NAME}_${TIMESTAMP}.diagnostic.log"
# Skip reporting on clean operator shutdowns
if [ "${EXIT_STATUS}" = "0" ] && [ "${EXIT_CODE}" = "0" ]; then
exit 0
fi
mkdir -p "${DUMP_DIR}"
cat <<EOF > "${DUMP_FILE}"
================================================================================
OPENCLAW POST-MORTEM RUNTIME AUDIT
Timestamp: ${TIMESTAMP}
Service: ${SERVICE_NAME} (${INSTANCE_NAME})
Exit Reason: Status=${EXIT_STATUS} Code=${EXIT_CODE}
================================================================================
[SYSTEMD SERVICE STATUS]
$(systemctl status "${SERVICE_NAME}" --no-pager -l || true)
[CGROUP PEAK RESOURCE FOOTPRINT]
Memory Current: $(cat /sys/fs/cgroup/system.slice/openclaw.service/memory.current 2>/dev/null || echo "N/A") bytes
Memory Peak: $(cat /sys/fs/cgroup/system.slice/openclaw.service/memory.peak 2>/dev/null || echo "N/A") bytes
Memory Events: $(cat /sys/fs/cgroup/system.slice/openclaw.service/memory.events 2>/dev/null || echo "N/A")
CPU Stat: $(cat /sys/fs/cgroup/system.slice/openclaw.service/cpu.stat 2>/dev/null || echo "N/A")
[LAST 50 LOG ENTRIES PRIOR TO TERMINATION]
$(journalctl -u "${SERVICE_NAME}" -n 50 --no-pager -o short-iso)
================================================================================
EOF
chmod 640 "${DUMP_FILE}"
logger -t openclaw-supervisor "Diagnostic dump recorded at ${DUMP_FILE} for exit state ${EXIT_STATUS}."
Ensure correct execution permissions:
chmod +x /usr/local/bin/openclaw-crash-reporter.sh
5. High-Throughput Journald Logging Hygiene and Fast Audits
Autonomous 24/7 processing generates high write-rates across standard input/output. Unbounded logging can rapidly fill local partitions, trigger journal rotation thrashing, and exhaust storage bandwidth.
To preserve NVMe IOPS for active database operations and LLM session state, configure a dedicated systemd-journald drop-in configuration for OpenClaw at /etc/systemd/journald.conf.d/01-openclaw.conf:
[Journal]
Storage=persistent
Compress=yes
RateLimitIntervalSec=30s
RateLimitBurst=10000
SystemMaxUse=2G
SystemKeepFree=4G
SystemMaxFileSize=128M
MaxRetentionSec=1month
Apply the journald settings without dropping the log stream:
systemctl kill --kill-who=main --signal=SIGUSR1 systemd-journald
With write rate-limiting and rotation boundaries enforced, utilize these deterministic commands for continuous auditing and immediate incident triage:
Live Tail of Errors, Warnings, and Circuit-Breaker Alerts
Filter out routine event loops and stream only actionable exceptions:
journalctl -u openclaw.service -f -p err..emerg -o cat
Correlating Watchdog Kills and Exit Codes
Locate every automated supervisor restart event over the previous 48 hours:
journalctl -u openclaw.service --since "48 hours ago" \
| grep -E "Watchdog timeout|Failed with result|Scheduled restart job"
Real-Time cgroup Resource Consumption
Inspect CPU usage, thread counts, and memory consumption live without launching high-overhead tools like top:
systemd-cgtop /system.slice/openclaw.service
This systems-level supervision architecture guarantees that when running OpenClaw on an enterprise KVM instance from tropic.host, unhandled application state exceptions are converted into controlled, sub-second restarts. By pairing cgroups v2 resource fencing with hardware-isolated execution, your autonomous stack maintains continuous operational continuity around the clock.
Frequently Asked Questions (FAQ)
Why run autonomous AI agents on a VPS instead of a local PC?
A KVM VPS provides 99.99% uptime, low-latency connectivity to AI APIs, a dedicated static IP, and safe sandboxed execution without consuming local PC resources 24/7.
What happens if an autonomous agent executes unsafe bash commands?
By isolating OpenClaw inside Docker containers with dropped root capabilities and non-root users, system hosts remain fully protected against rogue actions.