Review ID: 01339e510d66Generated: 2026-08-29T23:39:59.345Z
CHANGES REQUESTED
26
Raw Findings
14
High
11
Low
691/ 1000
ShipItClean Score · Progressing
4 of 108 Agents Deployed
DiamondPlatinumGoldSilverBronzeHR RoastyFree Baseline
2 Diamond · 1 Silver · 1 Bronze
laude-institute/headlong →
main @ 66e99b7
AIAI Threat Analysis
# Triage Report: laude-institute/headlong
REAL THREATS
Authentication & Authorization Bypass (Web Server)
[8] Remote Code Execution via Unauthenticated Self-Update The /api/update endpoint executes git pull --ff-only and restarts the server without any authentication. While gated by an environment flag (HEADLONG_WEB_SELF_UPDATE=1), once enabled it accepts requests from any localhost process. If the git remote has been compromised or an attacker gains any local access, they can trigger arbitrary code deployment and execution.
[9] Unauthenticated Identity Data Exfiltration /api/identities/{identity_id}/export serves complete identity archives (memories, trajectories, .env files with potential secrets) with zero authentication. Any process on localhost can enumerate and download all identity data.
[13] Denial of Service via Killall /api/killall terminates every Headlong process system-wide without authentication. On localhost-only deployments, any local user or compromised service can halt the entire platform.
[14] Environment Variable Injection PUT /api/identities/{identity_id}/env writes arbitrary key-value pairs to identity .env files without authentication. An attacker can inject ANTHROPIC_API_KEY or override MODEL settings to redirect LLM calls, exfiltrate data, or cause billing fraud.
State & Concurrency Issues
[3] State File Corruption on Crash state.py:21 writes the offset directly without atomic rename. A crash mid-write leaves a truncated or corrupted file. On restart, int(path.read_text()) can fail or load garbage, breaking message deduplication in the Telegram bot and potentially causing infinite loops or message loss.
[11] Race Condition in Identity Import Concurrent import requests share a temp file created with tempfile.mkstemp(). The function doesn't prevent two requests from writing to the same file descriptor simultaneously, leading to archive corruption and potential identity data loss or mixing.
[12] Race Condition in Identity Creation Concurrent POST /api/identities calls can mkdir and write to the same directory without locking, resulting in half-written configs, corrupted info.txt, or unusable identities.
Operational & Configuration Issues
[1] Memory Exhaustion in Telegram Dedup Cache The _seen_msg_ids dict grows unbounded during message floods. While pruned on each new message (retaining 300s window), an attacker sending thousands of messages per second can force multi-GB memory growth before the next prune cycle, potentially OOM-killing the bot.
[4] Silent Authentication Failure Defaulting missing ANTHROPIC_API_KEY to empty string causes silent auth failures that may be obscured by shell pipelines. While not directly exploitable, it creates operational blind spots that mask real attacks or misconfigurations.
ATTACK CHAINS
Chain 1: Localhost Takeover → Full System Compromise
1. Attacker gains any localhost access (e.g., via browser exploit, SSRF from another service, or shared development machine)
2. [14] Inject malicious API key into identity .env → redirect LLM calls to attacker-controlled endpoint → exfiltrate all conversation data
3. [9] Export all identities → extract secrets, conversation history, and PII
4. [8] If HEADLONG_WEB_SELF_UPDATE=1, pull malicious code and restart → persistent RCE
5. [13] Or simply /api/killall to deny service
Chain 2: Multi-Request Race → Identity Corruption
1. [12] Send concurrent identity creation requests with same name
2. [11] Follow immediately with concurrent import requests
3. Result: corrupted identity directory, mixed archive data, potential secret leakage across identities
DROPPED FINDINGS (False Positives / Low Impact)
[2] HTML injection in Telegram - The rationale is truncated but describes proper escaping via to_html(). Telegram's HTML mode is restricted to their safe subset; this is not XSS in a browser context. FALSE POSITIVE.
[5] Path traversal in identity import - The function signature shows archive: Path (typed), and without evidence of user-controlled input reaching this from an API endpoint, this is speculative. INSUFFICIENT EVIDENCE.
[6][7] SSRF and path traversal in openrouter.py - No code snippet provided, only generic scanner output. Line 53/58 likely reference OpenRouter API calls (fixed endpoint) or model cache paths. Without seeing actual user input flow, INSUFFICIENT EVIDENCE.
[10] Job download IDOR - UUIDs are cryptographically random (128-bit). While lack of auth is bad practice, guessing a UUID job_id has 2^-128 probability per attempt. LOW EXPLOITABILITY (demoted from HIGH).
[15-25] Shell quoting issues - Standard shellcheck warnings. While these can cause issues with filenames containing spaces, they're code quality issues, not security vulnerabilities in this context. ACCEPTED AS LOW.
[0] Insecure HTTP - INFO severity, single mention in Swift macOS app. Without context (could be localhost-only), and marked INFO by scanner. ACCEPTED AS INFO.
VERDICT
This system is NOT safe to deploy in any multi-tenant or network-accessible environment.
The web server effectively has no authentication layer while exposing:
• Complete data exfiltration (identities, conversation histories, secrets)
• Environment manipulation (API key injection, model override)
• Code execution (if self-update is enabled)
• Denial of service
The implicit security model assumes "localhost-only = trusted," which fails catastrophically on:
• Shared development machines
• Systems running other web services (SSRF pivot points)
• Any scenario where an attacker achieves initial foothold
Must fix before ANY deployment:
1. [8][9][13][14] - Implement authentication middleware for ALL /api/* endpoints. At minimum, require a bearer token; ideally integrate with existing identity system.
2. [3] - Use atomic write pattern (write to temp, rename) for state file.
3. [11][12] - Add file locking or identity-level mutexes for import/create operations.
Should fix for production:
4. [1] - Add hard limit to dedup cache size (e.g., max 10K entries, LRU eviction).
5. [4] - Fail fast with clear error when API key is missing.
Risk assessment: Currently CRITICAL for any deployment. With authentication and atomic writes: MEDIUM (standard operational risks remain).
26 raw scanner findings — 14 high · 11 low · 1 info
▶ Raw Scanner Output — 26 pre-cleanup findings
⚠ Pre-Cleanup Report
This is the raw, unprocessed output from all scanner agents before AI analysis. Do not use this to fix issues individually. Multiple agents attack from different angles and frequently report the same underlying vulnerability, resulting in significant duplication. Architectural issues also appear as many separate line-level findings when they require a single structural fix.

Use the Copy Fix Workflow button above to get the AI-cleaned workflow — it deduplicates findings, removes false positives, and provides actionable steps. This raw output is provided for transparency and audit purposes only.
HIGHUnbounded dedup cache grows without limit under sustained load
telegram/src/headlong_telegram/inbound.py:56
[AGENTS: Chaos]memory-exhaustion
The `_seen_msg_ids` dict is pruned only when a new message arrives (line 114-117). If an attacker sends a flood of messages, the dict grows to hold all of them within the 300-second window. With a high message rate (e.g., 1000 msg/sec), this could hold 300,000 entries, consuming significant memory. More critically, the pruning only happens when a message with a message_id arrives, so if the attacker sends messages with monotonically increasing IDs, the dict keeps growing. The same unbounded pattern exists in slack/state.py Deduper (bounded at 5000, but still large) and outbound.py RecentPosts.
Suggested Fix
Add a hard cap on the dict size and prune aggressively, or use a TTL-based cache.
HIGHHTML injection in Telegram messages via unescaped content
telegram/src/headlong_telegram/outbound.py:84
[AGENTS: Chaos]html-injection
The `to_html` function escapes prose text but the `_convert_code` function escapes code spans. However, the `_convert_prose` function converts markdown links to HTML anchors using `re.sub(r"\[([^\]]+)\]\(([^)\s]+)\)", r'<a href="\2">\1</a>', text)`. The URL in `\2` is NOT validated or escaped. A malicious agent reply containing `[text](javascript:alert(1))` or `[text](data:text/html,...)` would produce an HTML anchor with a dangerous href. While Telegram's HTML parser may sanitize some of this, the URL is attacker-controlled and could contain quote characters to break out of the href attribute. The same issue exists in slack/slackfmt.py line 18.
Suggested Fix
Validate URLs in markdown links (only allow http/https) and escape quote characters in the href.
HIGHNon-atomic offset write can corrupt state on crash
telegram/src/headlong_telegram/state.py:21
[AGENTS: Chaos]state-corruption
The offset is written directly to the state file without atomic write (no temp file + rename). If the process crashes mid-write, the file can be truncated or contain partial data. On restart, `int(path.read_text().strip())` will raise ValueError, which is caught and the offset resets to 0. This causes the bridge to replay ALL historical updates from the beginning, potentially re-delivering old messages to users. The same non-atomic write pattern exists in slack/state.py ActiveThreads._save (line 84) and telegram/allowlist.py _save (line 106).
Suggested Fix
Write to a temp file and atomically rename: `tmp = self._path.with_suffix('.tmp'); tmp.write_text(str(value)); tmp.replace(self._path)`
HIGHEmpty API key silently passed to agent, causing silent failures
terminal_bench2_eval/harbor_headlong_agent.py:206
[AGENTS: Chaos]secrets-management
If ANTHROPIC_API_KEY is not set in the harness environment, an empty string is passed to the agent. The agent will fail authentication, but the error may be swallowed by the `2>&1 | tee` pipeline and the run will appear to hang until timeout. This wastes the entire task budget (up to 1000 iterations) on silent auth failures. The code should fail fast at setup if the key is missing.
Suggested Fix
Validate the key exists before starting: `if not os.environ.get("ANTHROPIC_API_KEY"): raise RuntimeError("ANTHROPIC_API_KEY not set")`
HIGHUnsanitized archive path and name in identity import
web/src/headlong_web/control.py:420
[AGENTS: Chaos]path-traversal
OWASP A01:2021NIST AC-3
The `archive` path is passed directly to the identity CLI import command. If an attacker can control the archive path (e.g., via a crafted request), they could import from an arbitrary file path on the server, potentially reading sensitive files. Additionally, the `name` parameter is passed as `--name` without validation — a malicious name could cause the import to write to an unexpected location or overwrite existing identities. The archive itself is a tar.gz file that could contain path traversal entries (e.g., `../../etc/cron.d/evil`), which the identity CLI might extract outside the intended directory if it doesn't sanitize archive entries.
Suggested Fix
Validate the archive path is within an allowed temp directory, validate the name against a strict pattern, and ensure the identity CLI sanitizes tar entry paths during extraction.
HIGHSSRF — user-controlled URL in HTTP request
web/src/headlong_web/openrouter.py:53
[AGENTS: rules-engine]security
HTTP request with user-controlled URL in web/src/headlong_web/openrouter.py at line 53 enables SSRF attacks.
Suggested Fix
Validate URLs against an allowlist. Block private/internal IP ranges.
HIGHPath traversal — user input in file operation
web/src/headlong_web/openrouter.py:58
[AGENTS: rules-engine]security
File operation with user-controlled path in web/src/headlong_web/openrouter.py at line 58. Attacker can read arbitrary files via ../ sequences.
Suggested Fix
Validate and sanitize file paths. Use path.resolve() and verify the result is within the expected directory.
HIGHUnauthenticated self-update endpoint can pull attacker-controlled code and restart the server
web/src/headlong_web/server.py:235
[AGENTS: Chaos]rce
The POST /api/update endpoint runs `git pull --ff-only` on the code repo and then restarts the process. It is gated only by HEADLONG_WEB_SELF_UPDATE=1 and read_only. There is no authentication. If the repo's remote is attacker-influenced (or the git config is manipulated), an unauthenticated caller can trigger a pull of arbitrary code and a restart, achieving remote code execution on the server. Even without remote control, an attacker can repeatedly trigger pulls/restarts to cause a DoS. The endpoint is enabled on the demo deployment by default.
Suggested Fix
Require authentication and a signed/verified update source.
HIGHUnauthenticated identity export leaks full identity data including memories and trajectory
web/src/headlong_web/server.py:780
[AGENTS: Chaos]idor
The export endpoint builds a .shellm.tgz archive of the identity's entire directory (memories, trajectory, .env, etc.) and serves it as a download. There is no authentication. Any local process or network peer that can reach the web port can download any identity's full data, including the .env file which contains API keys (the export includes the whole identity dir). This is a direct data exfiltration vulnerability.
Suggested Fix
Require authentication for export endpoints, and exclude .env from exports or encrypt it.
HIGHExport job download endpoint leaks any identity's archive by job_id without auth
web/src/headlong_web/server.py:917
[AGENTS: Chaos]idor
The export job download endpoint serves the finished archive file by job_id. Job IDs are UUIDs (uuid.uuid4().hex), which are not guessable, but there is no authentication and no ownership check. If an attacker can observe or guess a job_id (e.g., from logs, referrer, or a shared browser), they can download any identity's full export, including memories, trajectory, and potentially secrets. The job list endpoint /api/identities/{identity_id}/export-jobs also returns job metadata without auth, and the download URL is predictable once the job_id is known. This is an IDOR/unauthorized data access issue.
Suggested Fix
Require authentication and tie job ownership to the requesting identity.
HIGHConcurrent identity imports race on the same temp file and can corrupt the archive
web/src/headlong_web/server.py:933
[AGENTS: Chaos]concurrency
The import endpoint creates a temp file with tempfile.mkstemp() and streams the request body into it, then calls control.identity_import() in a threadpool. If two requests arrive concurrently, each gets its own temp file, so no direct collision. However, control.identity_import() extracts the archive into the identities directory without a lock. Two concurrent imports of the same identity name can both extract, then both rename/move directories, causing one to fail with a partial/corrupted identity left behind. Also, the temp file is unlinked in the finally block while the threadpool task may still be reading it (run_in_threadpool returns a coroutine that is awaited, so the finally runs after the task completes — but if the task raises, the file is removed before the error is surfaced, losing the diagnostic). The bigger issue: no per-identity lock means concurrent imports of the same name can interleave directory creation and leave a half-written identity.
Suggested Fix
Serialize imports per identity name with a lock, and validate the archive is fully written and closed before handing it to the import routine.
HIGHConcurrent identity creation can race and produce a corrupted identity directory
web/src/headlong_web/server.py:968
[AGENTS: Chaos]concurrency
create_identity calls control.identity_new(root, body.name) with no lock. If two requests create the same identity name concurrently, both may mkdir the same directory, both write info.txt/activate, and one may overwrite the other's files mid-write, leaving a corrupted identity. Also, the name is validated with IDENTITY_NAME_RE, but the check happens in the route; control.identity_new may not re-validate, so a race between validation and use is possible. This is a state-corruption risk under concurrent access.
Suggested Fix
Add a per-identity creation lock and check for existing directory before creating.
HIGHUnauthenticated killall endpoint can terminate every Headlong process on the host
web/src/headlong_web/server.py:978
[AGENTS: Chaos]dos
The /api/killall endpoint calls control.killall() which shells out to headlong-killall and kills every Headlong process (dispatchers, thinker steps, the dash). The web server binds localhost by default but 0.0.0.0 in a container, and there is NO authentication on any endpoint. Any process on the host (or any container on the same network) can POST to /api/killall and stop the entire agent fleet. This is a trivial DoS with no auth requirement. The read_only flag only gates it when explicitly set; the default deployment has controls enabled.
Suggested Fix
Require authentication (token) for all mutating endpoints, especially killall.
HIGHUnauthenticated env write lets any local process inject API keys or override model settings
web/src/headlong_web/server.py:1021
[AGENTS: Chaos]auth
The PUT /api/identities/{identity_id}/env endpoint writes arbitrary key/value pairs into the identity's .env file. There is no authentication. An attacker who can reach the web port (localhost on a shared machine, or the container's published port) can set ANTHROPIC_API_KEY to their own key, redirecting the agent's LLM calls to their account, or set SHELLM_MODEL to a malicious value. This is a privilege escalation / data-theft vector: the attacker can also read the existing keys via GET /api/identities/{identity_id}/env (which returns redacted entries, but the redaction may not mask the full key — redacted_entry is used, but the value may still be partially visible). More critically, the attacker can overwrite the key with their own, causing the agent to send all its prompts (including sensitive conversation content) to the attacker's LLM endpoint.
Suggested Fix
Require authentication for env read/write endpoints, and never expose even redacted keys to unauthenticated callers.
LOWUnquoted Command Substitution In Command
deploy/migrate-units.sh:414
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
deploy/scripts/lib.sh:47
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
deploy/thinkers-death-alert.sh:82
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
deploy/thinkers-failure-alert.sh:41
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
deploy/thinkers-service.sh:63
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
deploy/update.sh:23
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
install.sh:188
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
status.sh:48
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
thinkers/_lib/common.sh:41
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
thinkers/retrieval/build-index.sh:19
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
LOWUnquoted Command Substitution In Command
uninstall.sh:87
[AGENTS: rules-engine]code_quality
The result of command substitution $(...) or `...`, if unquoted, is split on whitespace or other separators specified by the IFS variable. You should surround it with double quotes to avoid splitting the result.
INFOInsecure Http Request
macos/Shellm/main.swift:224
[AGENTS: rules-engine]security
The software transmits sensitive or security-critical data in cleartext in a communication channel that can be sniffed by unauthorized actors.
Suggested Fix
See CWE-319: Cleartext Transmission of Sensitive Information
Note: Fixing issues can create a domino effect — resolving one finding often surfaces new ones that were previously hidden. Multiple scan-and-fix cycles may be needed until you’re satisfied no further issues remain. How deep you go is your call.