Review ID: fc8c74a65957Generated: 2026-08-30T00:30:37.452Z
CHANGES REQUESTED
186
Raw Findings
16
Critical
90
High
21
Medium
26
Low
179/ 1000
ShipItClean Score · Critical Risk
8 of 108 Agents Deployed
DiamondPlatinumGoldSilverBronzeHR RoastyFree Baseline
2 Diamond · 6 Gold
volcengine/OpenViking →
main @ e8cedae
AIAI Threat Analysis
Loading AI analysis...
186 raw scanner findings — 16 critical · 90 high · 21 medium · 26 low · 33 info
▶ Raw Scanner Output — 186 pre-cleanup findings
⚠ Pre-Cleanup Report
This is the raw, unprocessed output from all scanner agents before AI analysis. Do not use this to fix issues individually. Multiple agents attack from different angles and frequently report the same underlying vulnerability, resulting in significant duplication. Architectural issues also appear as many separate line-level findings when they require a single structural fix.

Use the Copy Fix Workflow button above to get the AI-cleaned workflow — it deduplicates findings, removes false positives, and provides actionable steps. This raw output is provided for transparency and audit purposes only.
HIGHClaude Code invoked with dangerously-skip-permissions flag
benchmark/locomo/claudecode/eval.py:218
[AGENTS: Lockdown - Supply]Insecure Default, supply_chain
**Perspective 1:** The `claude` subprocess is invoked with the `--dangerously-skip-permissions` flag, which disables all permission checks. This means the LLM agent can execute arbitrary commands, read/write files, and access network resources without restriction. Combined with the API key in the environment, a compromised or malicious prompt could lead to data exfiltration or system compromise. This is a supply chain risk because the agent's behavior is not constrained. **Perspective 2:** The eval script passes the --dangerously-skip-permissions flag to the claude CLI, which disables all permission checks for file system and command execution. While this is a benchmark script, if the prompts or project directories contain untrusted content, the LLM could be tricked into executing arbitrary commands or reading sensitive files without any guardrails. This is an insecure default for the evaluation environment.
Suggested Fix
Use a more restrictive permission model or sandbox the execution environment to limit the impact of the skip-permissions flag.
HIGHWhatsApp client fetches latest Baileys version at runtime without integrity verification
bot/bridge/src/whatsapp.ts:49
[AGENTS: Supply]supply_chain
The WhatsApp bridge calls fetchLatestBaileysVersion() at runtime to determine which Baileys library version to use. This means the actual library version used is not pinned or verified — it depends on the latest version available from the registry at connect time. A compromised or malicious Baileys release could be pulled and executed without any integrity check, and the behavior of the bridge is non-deterministic across deployments. This is a dependency confusion / supply chain risk.
Suggested Fix
Pin the Baileys version explicitly in package.json and import the pinned version. Do not fetch the latest version at runtime. Verify the package integrity (e.g., via lockfile and checksum).
HIGHTruncated code in message_router_loop causes NameError and crashes the router
bot/demo/werewolf/werewolf_server.py:1433
[AGENTS: Chaos]correctness
The code at line 1433 is truncated (`state=stat`), which is a syntax error / NameError. This will crash the entire message router loop, stopping the game. This is a critical correctness bug.
Suggested Fix
Fix the truncated code to `state=state`.
HIGHMutable 'latest' image tag used by default
bot/deploy/docker/build-image.sh:17
[AGENTS: Supply]supply_chain
The build script defaults to the mutable `latest` tag for the container image. This makes it impossible to verify which exact code is running in production, and an attacker who can push a malicious image to the registry could overwrite the `latest` tag, causing the deployment to run untrusted code. The fix is to require an immutable, unique tag (e.g., git SHA) and avoid `latest`.
Suggested Fix
Require IMAGE_TAG to be set to a unique, immutable value (e.g., git SHA) and refuse to build with 'latest'.
HIGHRemote skill file download lacks integrity verification
bot/vikingbot/agent/remote_skills.py:1162
[AGENTS: Supply]supply_chain
Remote skill files are downloaded from the OpenViking server and written to the sandbox without verifying the file's SHA-256 checksum against the manifest. The manifest includes a sha256 field for each file, but the download path does not validate that the downloaded bytes match the expected checksum. An attacker who can tamper with the server response (e.g., MITM, compromised server) could inject malicious skill content that gets executed in the sandbox.
Suggested Fix
After downloading, compute the SHA-256 of the received bytes and compare against the manifest's sha256 field before writing to disk.
HIGHSSRF via unvalidated image URL in _url_to_base64
bot/vikingbot/agent/tools/image.py:134
[AGENTS: Sentinel]SSRF / Unvalidated URL
The `_url_to_base64` method fetches any URL supplied by the user/agent without validation. An attacker controlling the `base_image` or `mask` parameter (e.g., via a prompt injection or a malicious agent) can point it at internal services (e.g., http://169.254.169.254/latest/meta-data/, http://localhost:xxxx) and exfiltrate their content by having the response base64-encoded and returned to the caller. The URL is passed directly to `httpx.AsyncClient.get(url)` with no allowlist, scheme, or host validation.
Suggested Fix
Validate the URL against an allowlist of permitted hosts/schemes before fetching, and block private/loopback/link-local IP ranges. Alternatively, require the image to be provided as a data URI or sandbox-local path only.
HIGHPath traversal in openviking_export via unvalidated dest and file URI
bot/vikingbot/agent/tools/ov_file.py:1165
[AGENTS: Sentinel]input_validation
The `dest` parameter (user/LLM-controlled) and the file URI-derived local path are concatenated without sanitization and passed to `sandbox.write_file(relative, text)`. A malicious `dest` value such as `../../etc` or a viking URI containing `../` segments can escape the intended sandbox workspace directory and overwrite arbitrary files the sandbox user can write. Trace: tool parameter `dest`/`uri` -> `local_path_for_viking_uri` -> `relative` -> `sandbox.write_file`.
Suggested Fix
Validate `dest` against a safe pattern (e.g. `^[\w.-]+$`), reject `..` segments, and resolve the final path with `Path(...).resolve()` then verify it stays under the sandbox workspace root before writing.
HIGHSSRF via unvalidated URL in web_fetch tool
bot/vikingbot/agent/tools/web.py:112
[AGENTS: Chaos - Lockdown - Sentinel]input_validation, insecure_configuration, ssrf
**Perspective 1:** The `web_fetch` tool validates that the URL uses http/https and has a netloc, but does not restrict the hostname or IP. An agent (or a user controlling the agent's input) can fetch `http://169.254.169.254/`, `http://localhost:1933/`, or other internal endpoints, enabling SSRF to internal services, cloud metadata, or the OpenViking server itself. The URL enters from the tool invocation, which is driven by LLM/user input. **Perspective 2:** The `web_fetch` tool validates only that the URL uses http/https and has a netloc, but does not block private/internal IP addresses (e.g., 127.0.0.1, 169.254.169.254, 10.x.x.x). An attacker who can control the `url` parameter (e.g., via a prompt injection or a malicious agent) can make the bot fetch internal metadata endpoints or internal services, leaking sensitive data. This is a direct input-to-sink path: `url` -> `client.get`. **Perspective 3:** The web_fetch tool uses `httpx.AsyncClient` without explicitly configuring TLS certificate verification. While httpx verifies certificates by default, the code does not set `verify=True` explicitly, and there's no handling for certificate errors. If the default behavior is changed or if a custom client is passed, this could allow man-in-the-middle attacks.
Suggested Fix
Block private, loopback, and link-local IP ranges, and resolve DNS to verify the target is not internal before fetching.
HIGHChannel media handling reads arbitrary local files from user-supplied paths
bot/vikingbot/channels/base.py:245
[AGENTS: Lockdown]Arbitrary File Read
In _parse_data_uri, when a data_uri does not start with data:, http://, https://, or send://, the code treats it as a local file path and reads it with path_obj.read_bytes(). An attacker who can send a message to the bot can supply an absolute path such as /etc/passwd or /home/user/.ssh/id_rsa and the bot will read and return the file contents. The input originates from the chat message boundary and reaches a file-read sink with no path validation or restriction to a safe directory.
Suggested Fix
Restrict local file reads to a configured safe directory (e.g., the workspace) and reject absolute paths or paths containing '..'.
HIGHLocal file read in _parse_data_uri allows arbitrary file access
bot/vikingbot/channels/base.py:245
[AGENTS: Chaos - Lifeline - Sentinel]error_handling, input_validation, security
**Perspective 1:** The _parse_data_uri method, when given a non-URL, non-data URI string, treats it as a local file path and reads its bytes. An attacker who can control the 'media' list of an inbound message (e.g., via a crafted email or chat message) can read arbitrary files from the bot's filesystem. The input is the media URL/path from an inbound message, which is attacker-controlled. **Perspective 2:** In `_parse_data_uri`, when a data URI does not match `data:`, `http://`, `https://`, or `send://`, the code falls through to treat the input as a local file path and calls `path_obj.read_bytes()` on it. The `data_uri` originates from an untrusted chat message (e.g. a user sending a message with a media attachment or a URL-like string). An attacker who can send a message to the bot can supply an absolute path such as `/etc/passwd` or `../../etc/shadow` and the bot will read and return the file contents, leaking arbitrary local files. There is no validation that the path is within an allowed directory, no allowlist of extensions, and no size limit on the read. **Perspective 3:** When a data_uri is not a data:, http(s):, or send:// URI, the code treats it as a local file path and calls path_obj.read_bytes() directly. If the file does not exist or is unreadable, a raw OSError/FileNotFoundError propagates to the caller with no context about which URI was being parsed or that it was a local-path fallback. This makes debugging failures difficult and can crash the message handling pipeline with an unhelpful error.
Suggested Fix
Restrict local file resolution to a configured safe directory, reject absolute paths and `..` traversal, and enforce a maximum file size before reading.
HIGHSSRF — user-controlled URL in HTTP request
bot/vikingbot/channels/discord.py:228
[AGENTS: rules-engine]security
HTTP request with user-controlled URL in bot/vikingbot/channels/discord.py at line 228 enables SSRF attacks.
Suggested Fix
Validate URLs against an allowlist. Block private/internal IP ranges.
HIGHOpenViking API proxy forwards all requests without verifying the caller's identity
bot/vikingbot/channels/openapi.py:586
[AGENTS: Supply]supply_chain
The proxy route forwards all requests to the OpenViking API. The gateway token is verified, but the caller's OpenViking identity is not resolved or verified before forwarding. An attacker who has the gateway token (or can reach the gateway on localhost) can make arbitrary API calls to the OpenViking server, potentially accessing or modifying data they should not have access to.
Suggested Fix
Resolve and verify the caller's OpenViking identity before forwarding the request, and enforce authorization on the proxied endpoints.
HIGHTrusted auth mode accepts client-supplied account/user headers without verifying them against a trusted source
bot/vikingbot/channels/openapi.py:892
[AGENTS: Supply]supply_chain
In trusted auth mode, the gateway accepts X-OpenViking-Account and X-OpenViking-User headers from the client and uses them to build an OpenViking connection. While it does check the upstream health response for matching account/user, the headers themselves are client-controlled. An attacker who can reach the gateway can supply arbitrary account/user IDs, and if the upstream health check returns matching values (which it may for a valid API key), the attacker can impersonate any user. The API key is also optional in this mode if a configured key exists, but the account/user headers are not validated against the configured identity.
Suggested Fix
Validate the account/user headers against the configured identity or require the API key to be present and valid.
HIGHWhatsApp bridge connection uses unverified WebSocket and sends auth token in plaintext
bot/vikingbot/channels/whatsapp.py:43
[AGENTS: Chaos - Supply]error_handling, supply_chain
**Perspective 1:** The `WhatsAppChannel.start` method connects to the bridge URL using `websockets.connect` without explicitly verifying the TLS certificate or enforcing a secure `wss://` scheme. The `bridge_url` comes from the channel config, which may be attacker-influenced. The bridge auth token is sent over this connection without any integrity or confidentiality guarantee if the connection is not properly secured. An attacker who can intercept the connection (e.g., via a misconfigured proxy or DNS poisoning) can steal the token and impersonate the bot. **Perspective 2:** The `bridge_url` from config is passed directly to `websockets.connect`. If the config is malicious or misconfigured (e.g., `file://` or a non-websocket scheme), the connection will fail, but the error may not be clear. More importantly, if the bridge_url is attacker-controlled, it could be used to connect to an unintended WebSocket endpoint, potentially leaking data or causing unexpected behavior.
Suggested Fix
Enforce `wss://` for the bridge URL, explicitly verify TLS certificates, and consider mutual TLS or a signed token exchange.
HIGHEnvironment variable expansion in config file enables secret exfiltration via config
bot/vikingbot/config/loader.py:98
[AGENTS: Lockdown - Tripwire]config_security, supply-chain/config-injection
**Perspective 1:** `load_config` reads the ov.conf JSON and runs `os.path.expandvars` over the raw text before parsing. This expands `${VAR}` and `$VAR` from the process environment into the config. If an attacker can influence the config file content (e.g., a shared or writable config path, or a config imported from an untrusted source), they can inject `${OPENVIKING_ROOT_API_KEY}` or `${HOME}` style references and cause the bot to embed environment secrets (API keys, tokens) into the parsed config, which may then be logged, sent to the LLM provider, or persisted. Even without file write access, a malicious config supplied via `OPENVIKING_CONFIG_FILE` can read arbitrary environment variables into the bot's runtime config. **Perspective 2:** The config loader expands $VAR and ${VAR} in the JSON config text. If a variable is unset, it is left unchanged, but if a secret is embedded via env var and later logged or dumped, the expanded value could be exposed. More importantly, this allows the config file to reference environment variables, which is a legitimate pattern, but the expanded raw text is parsed and could be included in error messages if JSON parsing fails, potentially leaking secret values.
Suggested Fix
Do not expand environment variables in config files. If expansion is required, restrict it to an explicit allowlist of known config keys and validate the source of the config file.
HIGHHardcoded default Langfuse secret key
bot/vikingbot/config/schema.py:835
[AGENTS: Lockdown]Insecure Configuration
The LangfuseConfig model defines a hardcoded default secret_key value. If a deployment enables Langfuse without overriding this value, all instances share the same secret, allowing an attacker to forge or read observability data. This is a default credential that must be changed.
Suggested Fix
Remove the hardcoded default and require the secret_key to be explicitly configured, or generate a random value at startup when not provided.
HIGHSandbox defaults to direct backend with unrestricted execution
bot/vikingbot/config/schema.py:843
[AGENTS: Lockdown]Insecure Configuration
The sandbox backend defaults to DIRECT, which executes commands directly on the host without isolation. Combined with the default restrict_to_workspace=False, this allows the agent to run arbitrary commands and access the entire filesystem. This is an insecure default for a production agent.
Suggested Fix
Default to a sandboxed backend (e.g., SRT or Docker) and require explicit opt-in for the direct backend.
HIGHConsole config editor allows arbitrary config modification including secrets
bot/vikingbot/console/web_console.py:374
[AGENTS: Lockdown - Sentinel]Injection / Input Validation, Insecure Configuration
**Perspective 1:** The Gradio config editor exposes all configuration fields, including API keys, provider credentials, and sandbox settings, as editable form fields. Since the console is unauthenticated and binds to 0.0.0.0, any network attacker can read and modify the full bot configuration, including injecting malicious provider endpoints or exfiltrating secrets. **Perspective 2:** The Config tab's `save_config_fn` takes values from Gradio components (which are populated from the current config and editable by any user who can reach the console) and constructs a new `Config` object, then persists it via `save_config`. Since the console has no authentication and binds to 0.0.0.0, an attacker can modify any config field, including provider API keys, endpoints, and sandbox settings. This can lead to credential theft (by pointing endpoints to attacker-controlled servers), arbitrary command execution (via sandbox config), or denial of service. Trace: network attacker -> unauthenticated console -> Config tab -> `save_config_fn` -> `Config(**config_dict)` -> `save_config` -> config file overwritten.
Suggested Fix
Add authentication to the console and restrict config editing to admin users; avoid rendering secret fields as plaintext editable inputs.
HIGHStored XSS via unsanitized session content rendered as HTML
bot/vikingbot/console/web_console.py:432
[AGENTS: Lockdown - Sentinel]Insecure Configuration, XSS / Input Validation
**Perspective 1:** The `load_session` function reads session JSONL files from disk and injects the `content` field directly into an HTML string without any escaping. The resulting HTML is rendered in a `gr.HTML` component. Session content originates from user messages (attacker-controlled) and is persisted to disk. When an admin opens the Sessions tab and selects a session, any `<script>` or event-handler payload stored in a message is executed in the admin's browser context. Trace: user message -> session JSONL file -> `load_session` -> `gr.HTML` -> browser. This is a stored XSS with no sanitization at the sink. **Perspective 2:** Session message content is inserted directly into HTML strings without escaping. If a session contains attacker-controlled content (e.g., from a chat message), it will be rendered as HTML in the console, enabling stored XSS. Combined with the unauthenticated console, this could allow an attacker to execute arbitrary JavaScript in the context of any user viewing the console.
Suggested Fix
Escape all message content before embedding in HTML, e.g. use `html.escape(content)` or render via a safe markdown/text component instead of raw HTML.
HIGHStored XSS via raw JSONL line rendered as HTML
bot/vikingbot/console/web_console.py:443
[AGENTS: Sentinel]XSS / Input Validation
In the exception handler of `load_session`, when a line fails to parse as JSON, the raw line is inserted directly into an HTML string and rendered in `gr.HTML`. A malformed or attacker-crafted line in a session file containing HTML/script is executed in the admin's browser. Trace: session file line -> `load_session` exception path -> `gr.HTML` -> browser. No escaping is applied.
Suggested Fix
Escape the raw line with `html.escape(line)` before embedding in HTML.
HIGHConsole server binds to all interfaces without authentication
bot/vikingbot/console/web_console.py:561
[AGENTS: Lockdown - Sentinel]Authentication / Input Validation, Insecure Configuration
**Perspective 1:** The Vikingbot console server (Gradio app) binds to 0.0.0.0, exposing the full configuration editor, session viewer, and workspace file explorer to any network client. The console has no authentication, and the config editor allows modifying all bot settings including API keys and provider credentials. The workspace tab also allows reading any file under the workspace path. This is a direct unauthorized access path to sensitive data and configuration. **Perspective 2:** The Vikingbot console server (Gradio app) binds to `0.0.0.0` on port 18791 with no authentication. The console exposes the Config tab (which can read and write the bot's configuration, including API keys and provider credentials), the Sessions tab (which displays full conversation content), and the Workspace tab (which reads arbitrary files within the workspace). Any network-reachable attacker can access the console, modify configuration, and exfiltrate sensitive data. Trace: network attacker -> unauthenticated Gradio console -> config read/write, session content disclosure, file read. **Perspective 3:** The console server runs over plaintext HTTP with no TLS. Any credentials, session data, or configuration viewed or edited through the console can be intercepted by network observers. The health endpoint and OpenAPI router are also served over plaintext. **Perspective 4:** The console server runs with log_level="warning", suppressing access logs. This reduces visibility into who is accessing the unauthenticated console, hindering incident detection and forensics. **Perspective 5:** The Gradio/FastAPI console app does not set security headers such as Content-Security-Policy, X-Frame-Options, or X-Content-Type-Options. This increases the risk of clickjacking and other browser-based attacks against console users. **Perspective 6:** The console endpoints, including the config editor and session viewer, have no rate limiting. An attacker can brute-force or abuse the unauthenticated console without restriction.
Suggested Fix
Bind to 127.0.0.1 by default, add authentication (e.g., token or basic auth) before mounting the Gradio app, and restrict access to trusted networks.
HIGHFUSE mount exposes filesystem to all local users via allow_other
bot/vikingbot/openviking_mount/fuse_proxy.py:311
[AGENTS: Lockdown]configuration_security
The FUSE mount is created with allow_other=True, which makes the mounted filesystem readable and writable by every user on the host, not just the mounting user. The mount proxies to the OpenViking data directory (.original_files), so any local user can read, create, modify, or delete files in the OpenViking data store. This is a direct privilege escalation / data exposure path for multi-user hosts. The mount also runs with default_permissions=True but no allow_root, and the underlying OpenVikingMount.add_resource is invoked on every file write, so any local user can inject resources into the OpenViking index.
Suggested Fix
Remove allow_other=True (or gate it behind an explicit opt-in flag) so the mount is only accessible to the mounting user. If multi-user access is genuinely required, restrict it to a specific group via allow_other plus a mount option that limits access, and document the risk.
HIGHUnvalidated local file path read in VLM image preparation
openviking/models/vlm/backends/litellm_vlm.py:263
[AGENTS: Chaos - Lifeline - Sentinel]error_handling, input_validation, resource_management
**Perspective 1:** The `_prepare_image` method treats any string that does not start with 'http://' or 'https://' as a local filesystem path and reads it without validation. An attacker who can control the `images` argument (e.g., via a chat endpoint that accepts image paths) can read arbitrary files from the server's filesystem, including sensitive configuration files, API keys, or user data. The path is not resolved against a safe base directory, and no check is made to prevent path traversal (e.g., '../../etc/passwd'). This is a direct arbitrary file read vulnerability. **Perspective 2:** `_prepare_image` reads the entire image file into memory with no size limit. A user-supplied image path (e.g., via an API that accepts a file path) pointing to a multi-GB file will exhaust memory. This is reachable from any vision completion call that accepts a Path or local file string. **Perspective 3:** In _prepare_image, when the image is a local file path, the file is opened and read with no try/except. If the file is missing, unreadable, or a permission error occurs, the raw OSError propagates up through the VLM call stack without being wrapped or categorized. The caller receives a generic OSError with no indication that it was an image-loading failure, making it hard to distinguish from a model API error. There is no retry for transient I/O errors.
Suggested Fix
Validate that the path is within an allowed directory (e.g., using pathlib.Path.resolve() and is_relative_to()), reject paths containing '..' or absolute paths outside the allowed base, and consider restricting to a dedicated upload directory.
HIGHFeishu image relative path not validated against traversal before write
openviking/parse/accessors/feishu_accessor.py:840
[AGENTS: Sentinel]Input Validation
In `_materialize_drive_folder`, downloaded images are written via `image_path = markdown_path.parent / rel_path`, where `rel_path` comes from `_resolve_image_refs` and ultimately from Feishu document image tokens/URLs. If `rel_path` contains `../` or absolute path components, the write escapes the intended directory. The same pattern exists in `access()` at line 474. Input trace: Feishu document image reference -> `_resolve_image_refs` -> `rel_path` -> `Path` join -> `write_bytes`.
Suggested Fix
Validate each `rel_path` with `Path(rel_path).is_relative_to(markdown_path.parent)` and reject traversal before writing.
HIGHUnvalidated branch name injected into GitHub archive URL
openviking/parse/accessors/git_accessor.py:461
[AGENTS: Sentinel]input_validation
**Perspective 1:** The `branch` parameter originates from user-supplied kwargs (e.g., `branch` or `ref` in the access call). It is URL-encoded with `quote(..., safe="/")` but not validated against path traversal or control characters. A branch value like `../../etc/passwd` or `..%2F..` could be used to construct a URL that downloads an unintended archive or triggers SSRF-like behavior against GitHub. The encoded value is directly interpolated into `zip_url = f"https://github.com/{owner}/{repo_slug}/archive/{ref}.zip"`, and the response is written to disk and extracted. This allows an attacker to control the downloaded content and potentially write arbitrary files into the extraction directory. **Perspective 2:** The `branch` parameter is URL-encoded with `quote(..., safe="/")`, but control characters (e.g., `\n`, `\r`) are not explicitly rejected. While URL encoding neutralizes most injection, the encoded value is still used in a URL that is fetched and extracted. Control characters could potentially cause issues in the HTTP request or in the extraction path.
Suggested Fix
Validate the branch/ref against a strict allowlist (e.g., `^[A-Za-z0-9._/-]+$`) and reject any value containing path separators, `..`, or control characters before constructing the URL.
HIGHPath traversal — user input in file operation
openviking/parse/accessors/git_accessor.py:478
[AGENTS: rules-engine]security
File operation with user-controlled path in openviking/parse/accessors/git_accessor.py at line 478. Attacker can read arbitrary files via ../ sequences.
Suggested Fix
Validate and sanitize file paths. Use path.resolve() and verify the result is within the expected directory.
HIGHUnvalidated branch name injected into GitLab archive URL
openviking/parse/accessors/git_accessor.py:558
[AGENTS: Sentinel - Supply]input_validation, supply_chain
**Perspective 1:** The `branch` parameter (user-supplied via kwargs) is directly interpolated into the GitLab archive URL without any encoding or validation. A branch value containing `/`, `..`, or URL metacharacters can alter the URL path, potentially causing the download of an unintended archive or enabling path traversal in the resulting extraction. The response is written to disk and extracted, so an attacker controlling the branch can influence what files are written into the local filesystem. **Perspective 2:** The GitLab archive URL is built from a branch name or 'HEAD' (mutable reference). The downloaded content is not pinned to a specific commit, so the code being ingested can change between runs. This breaks build reproducibility and provenance tracking, and the downloaded archive is not verified against a known-good hash. **Perspective 3:** The `branch` parameter is directly interpolated into the GitLab archive URL without any encoding or validation. Control characters (e.g., `\n`, `\r`) are not rejected. While the URL is fetched via `urllib`, control characters could potentially cause issues in the HTTP request or in the extraction path.
Suggested Fix
Validate the branch/ref against a strict allowlist (e.g., `^[A-Za-z0-9._/-]+$`) and reject any value containing path separators, `..`, or control characters before constructing the URL.
HIGHHTTP accessor performs unvalidated SSRF requests to arbitrary URLs
openviking/parse/accessors/http_accessor.py:210
[AGENTS: Chaos - Sentinel - Tripwire]SSRF, dependency_scanner, ssrf
**Perspective 1:** The HTTPAccessor and URLTypeDetector accept arbitrary URLs from user input (e.g., a user-provided URL to ingest) and issue HEAD/GET requests to them. The request_validator is only applied when event_hooks are built, but the detect() method's HEAD request at line 210 and the download GET at line 581 can be called with request_validator=None, allowing requests to internal network addresses (e.g., 169.254.169.254, localhost, internal services). This enables SSRF attacks where an attacker can probe or access internal resources. The URL is attacker-controlled via the source parameter, and there is no host allowlist or scheme restriction beyond http/https. **Perspective 2:** The HTTPAccessor accepts arbitrary URLs from user input (e.g., via the 'viking add' command or API) and performs HEAD/GET requests without validating the target host against an allowlist. An attacker can supply a URL pointing to internal services (e.g., http://169.254.169.254/ or http://localhost:8080/admin) to probe or interact with internal network resources. The request_validator is optional and not enforced here. Trace: user-supplied URL -> _download_url -> _url_detector.detect -> client.head(url). **Perspective 3:** The HTTP accessor uses follow_redirects=True for both HEAD and GET requests. If a user-supplied URL redirects to an internal address (e.g., http://169.254.169.254/ or http://localhost:PORT/), the client will follow it. The request_validator may not be applied to redirected requests, and the network guard hooks only validate the initial request. This can enable SSRF to internal services.
Suggested Fix
Enforce a mandatory network guard (SSRF protection) that validates the target host against private/internal IP ranges and a configurable allowlist before any HTTP request is made, and reject requests when request_validator is not provided.
HIGHHTTP accessor downloads arbitrary URLs without SSRF validation
openviking/parse/accessors/http_accessor.py:581
[AGENTS: Chaos - Sentinel - Tripwire]SSRF, dependency_scanner, ssrf
**Perspective 1:** The _download_url method issues a GET request to an attacker-controlled URL. The request_validator is optional and only applied if provided; when absent, the request proceeds without any SSRF protection. An attacker can supply a URL pointing to internal services (e.g., http://169.254.169.254/latest/meta-data/, http://localhost:port/) and the server will fetch and return the content, enabling SSRF and potential data exfiltration. The URL is derived from user input (source parameter) and is not restricted to public hosts. **Perspective 2:** The HTTPAccessor performs a GET request to a user-supplied URL without validating the target host. This enables SSRF attacks to access internal services, cloud metadata endpoints, or other restricted resources. The request_validator is optional and not enforced here. Trace: user-supplied URL -> _download_url -> client.get(url). **Perspective 3:** The GET request in _download_url uses follow_redirects=True. A user-supplied URL that redirects to an internal service (e.g., cloud metadata endpoint) will be followed, and the downloaded content is written to disk and potentially parsed. The request_validator is only applied to the initial request, not to redirect targets.
Suggested Fix
Enforce a network guard (request_validator) that validates the resolved IP/host against an allowlist before making any request, and reject private/loopback/link-local addresses unless explicitly permitted.
HIGHSSRF — user-controlled URL in HTTP request
openviking/parse/accessors/web_feed_accessor.py:599
[AGENTS: rules-engine]security
HTTP request with user-controlled URL in openviking/parse/accessors/web_feed_accessor.py at line 599 enables SSRF attacks.
Suggested Fix
Validate URLs against an allowlist. Block private/internal IP ranges.
HIGHSSRF via unvalidated page URLs from sitemap/feed
openviking/parse/accessors/web_feed_accessor.py:630
[AGENTS: Sentinel]SSRF
OWASP A10:2021NIST SC-7
The WebFeedAccessor fetches URLs extracted from a user-supplied sitemap or feed. `_filter_entries` only checks `same_host_only` (default True) against the feed's netloc, but `_apply_robots` and `_mirror` fetch each `entry.url` without validating the resolved host against an allowlist or the network guard. A malicious sitemap can list internal URLs (e.g. http://169.254.169.254/...) that the accessor will fetch, enabling SSRF. The `request_validator` is only applied to the initial feed fetch via event hooks, not to per-page fetches. Trace: user-supplied feed URL -> `_collect` -> `_mirror` -> `client.get(entry.url)`.
Suggested Fix
Apply the same request validation hooks to per-page fetches, and enforce an explicit host allowlist/deny-list (e.g. block link-local and metadata IPs) for all mirrored URLs.
HIGHVideo parser reads entire file into memory without size limit
openviking/parse/parsers/media/video.py:83
[AGENTS: Chaos - Lifeline - Sentinel - Supply]error_handling, input_validation, resource_exhaustion, supply_chain
**Perspective 1:** The video parser reads the entire video file into memory via `file_path.read_bytes()` with no size limit. An attacker can upload a multi-GB video file, causing the server to exhaust memory and crash (OOM). The file is also written to viking_fs as `video_bytes`, doubling memory usage. This is a direct DoS vector via the upload/parse path. **Perspective 2:** The VideoParser reads the entire video file into memory with file_path.read_bytes() before any size limit is applied. An attacker who can upload or reference a large video file can cause the server to exhaust memory. The file size is never validated against a configured limit, and the magic-byte check happens only after the full file is loaded. Trace: untrusted video file path -> read_bytes() -> memory exhaustion. **Perspective 3:** The VideoParser reads the entire video file into memory without verifying its integrity (e.g., hash check against a known value). The file content is trusted as-is and written to storage. An attacker who can tamper with the source file could inject malicious content that is then persisted and served to users. **Perspective 4:** In `VideoParser.parse`, `file_path.read_bytes()` can raise an `OSError` (e.g., permission denied, I/O error) that propagates without being wrapped with context about the resource being parsed. This loses error categorization and makes debugging harder.
Suggested Fix
Stream the file in chunks, enforce a configurable maximum file size, and write to viking_fs via a streaming API instead of loading the whole file into memory.
HIGHUnbounded ZIP download buffers entire response in memory
openviking/parse/understanding_api.py:665
[AGENTS: Chaos - Lifeline - Sentinel]input_validation, missing_retry, resource_exhaustion
**Perspective 1:** The UnderstandingAPI downloads a ZIP file from a URL returned by the parser API (zip_url). The entire response is buffered into memory via rsp.content with no size limit. An attacker who controls the parser API response (or a malicious parser API server) can return a multi-GB ZIP, exhausting server memory. This is reachable from add_resource with a URL that routes to the UnderstandingAPI. The zip_url is attacker-influenced (it comes from the parser API response). **Perspective 2:** The `_download_zip` method fetches the `zip_url` returned by the Understanding API and writes the entire response body (`rsp.content`) to a temp file with no size limit. The `zip_url` originates from the remote Understanding API response, which is attacker-influenced when a user submits a URL for parsing. A malicious or compromised Understanding API (or a redirect target) can return an arbitrarily large body, exhausting local disk. This is a memory/disk exhaustion vector with no validation on response size. **Perspective 3:** The `_download_zip` method makes a single HTTP GET request with no retry logic. Transient network errors or 5xx responses will cause the entire parse operation to fail. The zip_url is typically a presigned URL that may be valid for a limited time, so a quick retry would likely succeed.
Suggested Fix
Stream the response and enforce a maximum byte count (e.g., read in chunks and abort once a configured limit is exceeded), and validate the Content-Length header before writing.
HIGHZIP bomb: unbounded extraction and per-file read into memory
openviking/parse/understanding_api.py:673
[AGENTS: Chaos]resource_exhaustion
OWASP A04:2021NIST SC-5
safe_extract_zip extracts the ZIP to a temp dir, but there is no limit on total extracted size or number of files. A malicious ZIP (zip bomb) can expand to fill the disk. Additionally, each file is read entirely into memory via child.read_bytes() before writing to VikingFS, so a single large file inside the ZIP can exhaust memory. The resource_name is derived from the ZIP content (archive_root) or user-supplied source_name, and is interpolated into the URI without sanitization, allowing path traversal via '..' in resource_name (e.g. if source_name is '../../etc/passwd').
Suggested Fix
Enforce a maximum total extracted size and per-file size, and sanitize resource_name to prevent path traversal.
HIGHPath traversal in temp_doc_uri via unsanitized resource_name
openviking/parse/understanding_api.py:679
[AGENTS: Chaos]path_traversal
OWASP A01:2021NIST AC-3
resource_name is derived from display_name (user-supplied source_name) or inferred from the URL path or ZIP archive root. It is interpolated directly into the VikingFS URI without sanitization. A user can pass source_name='../../etc' or a URL whose path contains '..', causing the mkdir/write to escape the intended temp directory and write to arbitrary VikingFS paths. This is reachable from add_resource with a URL and source_name.
Suggested Fix
Sanitize resource_name to a single path component (reject '/', '\', '..', null bytes).
HIGHNew-format API key identity is directly decodable and not integrity-protected
openviking/server/api_keys/new.py:143
[AGENTS: Chaos - Lockdown - Supply]authentication, authentication_bypass, supply_chain
**Perspective 1:** The new API key format is base64url(account_id).base64url(user_id).base64url(secret). The account_id and user_id segments are plaintext base64, so anyone who sees a key can decode the identity. More critically, the identity segments are not cryptographically bound to the secret. An attacker who obtains a valid key for one user can modify the account_id/user_id segments (re-encode a different identity) and, if the secret segment still matches a stored key for that identity, could impersonate another user. The resolve() path decodes the identity directly from the key and only verifies the secret against the stored key for the decoded identity, so a key with a valid secret but tampered identity would fail unless the secret also matches the target user. However, the lack of an HMAC or signature over the identity segments means the key format provides no integrity protection against substitution if a secret is reused or leaked across users. **Perspective 2:** In resolve(), for a new-format key, account_id and user_id are decoded directly from the client-supplied key string (parse_api_key), then the code checks whether that account/user exists and whether the secret matches. The identity is not independently derived from a server-side lookup keyed by a random token; it is entirely attacker-influenced. While the secret comparison provides some protection, this design means any leaked or brute-forced secret for one user can be replayed with a different account_id/user_id if the same secret value is reused across users (e.g., when a deterministic seed is used). This weakens the separation between identity and credential. **Perspective 3:** In resolve(), for a new-format key, the code decodes account_id and user_id from the key itself, looks up the stored user, and then compares the full api_key string against stored_key_or_hash using hmac.compare_digest. When api_key_hashing_enabled is False (the default per the constructor docstring), stored_key_or_hash is the plaintext full key, so this is correct. However, when hashing is enabled, the code checks `user_info.get("key_prefix", "").startswith("$argon2") or stored_key_or_hash.startswith("$argon2")` and then calls _verify_api_key(api_key, stored_key_or_hash). If an attacker supplies a key whose decoded account_id/user_id match a real user but whose secret segment is empty or malformed, the Argon2 verify will fail and fall through to legacy — but if the stored key is plaintext and the attacker can guess the secret segment (e.g. because it is short or derived from a weak seed), the compare succeeds. The real risk is that the new format leaks the identity in the key itself, so any leak of a key fragment (e.g. in logs that redact only the secret) reveals the account/user, and the secret is only 256 bits of hex — acceptable, but the identity disclosure combined with weak-seed derivation (see above) makes forgery practical.
Suggested Fix
Include an HMAC or signature over the account_id and user_id segments using a server-side secret, and verify it during resolve() before trusting the decoded identity.
HIGHPlaintext API key comparison enables timing-safe but unauthenticated identity spoofing when hashing is disabled
openviking/server/api_keys/new.py:167
[AGENTS: Lockdown - Supply]authentication_bypass, supply_chain
**Perspective 1:** When api_key_hashing_enabled is False (the default), the stored key is the full plaintext API key. resolve() compares the client-supplied key against the stored plaintext. If an attacker obtains the users JSON (e.g., via file read, backup, or path traversal), they can directly impersonate any user. Combined with the default of hashing disabled, this is a direct credential-theft path. The code does use hmac.compare_digest, which is timing-safe, but the underlying storage is plaintext. **Perspective 2:** When api_key_hashing_enabled is False (the default), the full API key is stored in plaintext in the users JSON file (stored_key = key). This means the entire credential is persisted at rest without any hashing, protected only by file-level AES encryption. If the storage layer is compromised or the file is exfiltrated, all API keys are directly readable, allowing full account takeover. The docstring even states 'Default: false - rely on file-level AES encryption for protection', which is a defense-in-depth gap. API keys should be stored as salted hashes (e.g., Argon2id) by default.
Suggested Fix
Enable Argon2id hashing for API keys by default, and never store plaintext keys at rest.
HIGHcreate_account allows caller-supplied seed for deterministic admin key
openviking/server/api_keys/new.py:199
[AGENTS: Chaos - Lockdown - Sentinel - Supply]access-control, input_validation, insecure_defaults, supply_chain
**Perspective 1:** create_account() accepts a caller-supplied seed and passes it to generate_api_key(), producing a deterministic admin API key. An attacker who can invoke create_account with a chosen seed can predict the admin key for a new account, gaining full admin access to that account. This is a self-service privilege escalation / authentication bypass. Input: create_account request with seed -> generate_api_key() -> predictable admin key. **Perspective 2:** create_account(account_id, admin_user_id, seed) passes the caller's `seed` directly to generate_api_key. Any caller who can invoke this method (e.g. an unauthenticated account-creation endpoint, or a lower-privilege user who can create sub-accounts) can choose a seed and compute the exact admin API key offline for any account_id/user_id. Combined with the new format's plaintext identity segments, this lets an attacker mint a valid admin key for an account they do not own, achieving full privilege escalation. The seed parameter is attacker-controlled input reaching a cryptographic key-derivation sink. **Perspective 3:** create_account passes the optional seed directly to generate_api_key. If a caller (e.g., an admin provisioning script or an API consumer) passes a low-entropy or predictable seed, the admin key is predictable, allowing an attacker to forge a root-level admin credential. The API does not enforce entropy requirements on the seed. **Perspective 4:** create_account() accepts a 'seed' parameter that is passed directly to generate_api_key(). If this seed is attacker-controlled (e.g., from an API request or CLI argument), the attacker can choose a known seed and predict the resulting admin API key, gaining full administrative access to the newly created account. Even if the seed is only from trusted internal callers, the ability to supply a deterministic seed undermines the randomness guarantee of the credential. The seed parameter should be removed from the public API surface or restricted to a securely managed server-side secret.
Suggested Fix
Remove the seed parameter from create_account() and register_user(), or require it to be a high-entropy server-side secret never exposed to callers.
HIGHdelete_account has no authorization check; any caller can delete any account
openviking/server/api_keys/new.py:261
[AGENTS: Chaos]access-control
delete_account delegates to self._legacy.delete_account(account_id) with no verification that the caller is root or an admin of the target account. The legacy implementation (already flagged) has no authorization gate, so any code path that reaches NewAPIKeyManager.delete_account — e.g. an admin API endpoint that does not itself verify the caller's role — lets an attacker delete any account, destroying all its users and keys. The new manager adds no defense, so the vulnerability is inherited and reachable through the new-format key path.
Suggested Fix
Add an authorization check in delete_account that verifies the caller is root or an admin of the target account.
HIGHregister_user allows caller-supplied seed for deterministic user key
openviking/server/api_keys/new.py:291
[AGENTS: Chaos - Lockdown - Sentinel - Supply]access-control, input_validation, insecure_defaults, supply_chain
**Perspective 1:** register_user() accepts a caller-supplied seed and derives the user's API key deterministically. An attacker who can call register_user with a chosen seed can predict the resulting user key, allowing them to authenticate as that user without knowing a stored secret. Combined with the ability to choose user_id, this enables impersonation of any user in an account. Input: register_user request with seed and user_id -> generate_api_key() -> predictable key. **Perspective 2:** register_user(account_id, user_id, role, seed) passes the caller's `seed` to generate_api_key. If the register_user endpoint is reachable by an authenticated user (or unauthenticated in some deployments), the caller can choose a seed and compute the exact API key for any user_id in an account they can name, then use that key to impersonate the user. The new format base64-encodes account_id and user_id, so the attacker only needs to guess/choose the seed to forge the secret. This is a direct credential-forgery / privilege-escalation path. **Perspective 3:** register_user passes the optional seed to generate_api_key. A predictable seed yields a predictable user key, enabling account takeover. The seed is exposed as a public parameter without entropy validation. **Perspective 4:** register_user() accepts a 'seed' parameter passed to generate_api_key(). If an attacker can control this seed (e.g., via an API endpoint that calls register_user), they can predict the API key issued to a new user, then use it to impersonate that user. Even with trusted callers, deterministic key derivation from a caller-supplied seed weakens credential security. The seed should be removed or restricted to a server-side secret.
Suggested Fix
Remove the seed parameter from register_user, or derive the secret from a server-side master key that callers cannot influence.
HIGHregenerate_key has no authorization check; any caller can rotate any user's key
openviking/server/api_keys/new.py:375
[AGENTS: Chaos]access-control
regenerate_key(account_id, user_id, seed) checks only that the account and user exist, and that the user is not mid-deletion. It does not verify that the caller is root or an admin of the target account, or that the caller is the user themselves. Any code path that reaches NewAPIKeyManager.regenerate_key — e.g. an admin API endpoint that does not itself verify the caller's role — lets an attacker rotate any user's key, locking them out and obtaining the new key (which is returned). The new manager adds no authorization gate, so the vulnerability is inherited from the legacy design and reachable through the new-format key path.
Suggested Fix
Add an authorization check in regenerate_key that verifies the caller is root, an admin of the target account, or the user themselves.
HIGHregenerate_key allows caller-supplied seed for deterministic key
openviking/server/api_keys/new.py:417
[AGENTS: Chaos - Lockdown - Sentinel - Supply]access-control, input_validation, insecure_defaults, supply_chain
**Perspective 1:** regenerate_key() accepts a caller-supplied seed and derives the new key deterministically. An attacker who can trigger a key regeneration with a chosen seed can predict the new key for any user, enabling authentication bypass. Input: regenerate_key request with seed -> generate_api_key() -> predictable key. **Perspective 2:** regenerate_key(account_id, user_id, seed) passes the caller's `seed` to generate_api_key. If the regenerate endpoint is reachable by a user who can name an account_id/user_id (e.g. an admin or, in some flows, a self-service user), the caller can choose a seed and compute the exact new key offline, then use it to impersonate the target user. The new format's plaintext identity segments make this trivial once the seed is chosen. This is a direct credential-forgery / privilege-escalation path. **Perspective 3:** regenerate_key passes the optional seed to generate_api_key. If a caller supplies a predictable seed, the regenerated key is predictable, allowing an attacker to take over the account after a rotation. The API does not validate seed entropy. **Perspective 4:** regenerate_key() accepts a 'seed' parameter passed to generate_api_key(). If an attacker can control this seed, they can predict the new API key issued to a user after rotation, defeating the purpose of key rotation and allowing continued unauthorized access. The seed should be removed or restricted to a server-side secret.
Suggested Fix
Remove the seed parameter from regenerate_key, or require it to be a server-side secret that callers cannot influence.
HIGHset_role has no authorization check; any caller can escalate privileges
openviking/server/api_keys/new.py:452
[AGENTS: Chaos]access-control
set_role delegates directly to self._legacy.set_role(account_id, user_id, role) with no check that the caller is an admin or root. The legacy implementation (already flagged) has no authorization gate, so any code path that reaches NewAPIKeyManager.set_role — e.g. an admin API endpoint that does not itself verify the caller's role — lets an attacker escalate any user (including themselves) to admin. The new manager adds no defense, so the vulnerability is inherited and reachable through the new-format key path.
Suggested Fix
Add an authorization check in set_role that verifies the caller is root or an admin of the target account before allowing a role change.
HIGHDev auth plugin grants unauthenticated ROOT access when enabled
openviking/server/auth/plugins/dev.py:40
[AGENTS: Lockdown]configuration_security
The DevAuthPlugin.resolve_identity returns Role.ROOT for every request without any authentication, and accepts arbitrary x_openviking_account/x_openviking_user headers to impersonate any account/user. While validate_config() attempts to restrict this to localhost, the check only inspects config.host at startup. If the server is bound to 0.0.0.0 or behind a reverse proxy that forwards to localhost, or if the config is changed at runtime, this exposes an unauthenticated ROOT endpoint. An attacker who can reach the server can access all data and perform any action. The validate_config() gate is a defense-in-depth check that can be bypassed by misconfiguration, and the plugin itself is the insecure default when auth_mode='dev' is set.
Suggested Fix
Remove the dev plugin from production builds, or require an explicit opt-in flag AND verify the actual bound interface at request time. Never trust x_openviking_account/x_openviking_user headers without authentication.
HIGHReDoS via user-supplied regex pattern in grep endpoint
openviking/server/routers/search.py:528
[AGENTS: Sentinel]ReDoS
The /api/v1/search/grep endpoint accepts a user-supplied 'pattern' field and passes it directly to service.fs.grep without any validation. The grep implementation (as noted in prior findings) uses the pattern as a regex, allowing catastrophic backtracking (ReDoS) attacks. An attacker can supply a malicious regex pattern (e.g., '(a+)+$') that causes exponential CPU consumption, leading to denial of service. Trace: HTTP request -> GrepRequest.pattern -> service.fs.grep -> regex engine.
Suggested Fix
Validate the pattern against a safe regex allowlist, enforce a maximum pattern length, and use a timeout or non-backtracking regex engine.
HIGHWebDAV path normalization allows percent-encoded traversal bypass
openviking/server/routers/webdav.py:46
[AGENTS: Sentinel]input_validation
The WebDAV router decodes the URL path with unquote() before checking for '..' segments. However, the check happens after a single decode pass. An attacker can double-encode the traversal payload (e.g. %252e%252e%252f) so that unquote() yields '%2e%2e%2f' which does not match the literal '.'/'..' check, and the path is then passed to service.fs operations. A second decode happens downstream in the filesystem layer, turning it into a real '..' traversal that escapes the viking://resources root. Trace: HTTP request path -> _normalized_resource_path -> service.fs.rm/mv/write_file with the traversal path.
Suggested Fix
Decode repeatedly until stable (or reject any '%' after decoding), and validate the final decoded path has no '..' segments before use.
HIGHWebDAV MOVE Destination header allows cross-host path traversal
openviking/server/routers/webdav.py:164
[AGENTS: Sentinel]input_validation
The MOVE handler parses the Destination header with urlparse. If the header contains an absolute URL with a scheme (e.g. http://attacker.com/../../etc/passwd), parsed.path is used directly. The check only verifies the path starts with /webdav/resources, but the host portion is ignored, so an attacker can supply a Destination pointing to an arbitrary host while the path still passes the prefix check. Combined with the double-decode issue, this can move resources outside the intended root or to unintended locations. Trace: HTTP Destination header -> _destination_path -> service.fs.mv.
Suggested Fix
Reject Destination headers with a scheme/host that does not match the request's own host, and normalize/validate the final path after decoding.
HIGHUpload tokens are too short and exposed in URL query strings
openviking/server/upload_token_store.py:34
[AGENTS: Lockdown]configuration
Upload tokens are only 6 characters (base62), providing ~35.7 bits of entropy. They are exposed in URL query strings (?token=), making them visible in server logs, browser history, and referrer headers. An attacker who obtains or guesses a token can upload files as another user and trigger resource ingestion with the victim's identity and parameters. The token carries account_id, user_id, and business parameters, making this a privilege escalation vector.
Suggested Fix
Increase token length to at least 16 characters, use a one-time token that is consumed on first use, and avoid exposing tokens in URL query strings (use headers or POST body instead).
HIGHSession owner validation is inverted, allowing invalid user IDs to be used as migration targets
openviking/service/legacy_migration.py:361
[AGENTS: Chaos - Sentinel]input_validation, logic_error
**Perspective 1:** In `_plan_one_session`, the condition `if validate_user_id(owner):` treats a *valid* user ID as an error and a *invalid* one as acceptable. `validate_user_id` returns an error string when the ID is invalid and None when valid. The code appends an error to the plan when validation *succeeds* (returns None, which is falsy), and proceeds to create migration operations for invalid owner IDs. This allows a legacy session with a malicious or malformed `owner_user_id` (e.g. containing path traversal `../` or null bytes) to be copied into `user/{owner}/sessions/...`, potentially writing outside the intended user namespace or into another user's directory, causing cross-tenant data corruption or unauthorized data access. **Perspective 2:** The code checks `if validate_user_id(owner):` and treats a truthy return as an error, but `validate_user_id` typically returns an error string (truthy) on invalid input. If the function returns `None` or empty string on valid input, the logic is correct; however, if it returns a boolean `True` for valid input, this would incorrectly flag valid owners. This needs verification against the actual `validate_user_id` contract.
Suggested Fix
Invert the condition: `if not validate_user_id(owner):` so invalid IDs are rejected and valid ones proceed.
HIGHPrune orphan delete uses record owner_user_id to construct a delete context without verifying the caller can act for that owner
openviking/service/reindex_executor.py:920
[AGENTS: Chaos]access_control
In `_delete_ctx_for_prune_record`, when a record has an `owner_user_id` different from the caller's, the code constructs a new `RequestContext` with `UserIdentifier(ctx.account_id, str(owner))` and the caller's `role`. This context is then used to call `vikingdb.delete(ids, ctx=delete_ctx)`. If the caller is an ADMIN (or the role is preserved), the delete is performed under the owner's identity, bypassing any per-user authorization checks that would normally apply. An admin reindexing a namespace could delete vectors belonging to other users without verifying the caller has manage permission on those specific records. The `_is_orphan_vector_record` check reads the source file with `owner_ctx`, but the delete itself uses the same constructed context, so a malicious admin could prune (delete) another user's vectors by crafting a URI that matches their records. This is reachable from the admin prune_orphans endpoint.
Suggested Fix
Verify the caller has MANAGE permission on the record's URI (via ACL resolution) before constructing the owner delete context, or require ROOT role for cross-owner deletes.
HIGHArbitrary file deletion via unvalidated task path in delete
openviking/service/task_store.py:95
[AGENTS: Sentinel]input_validation
`PersistentTaskStore.delete` calls `_task_path` with user-controlled `account_id`, `user_id`, and `task_id`, then passes it to `agfs.rm(..., force=True)`. If any of these values contain path traversal sequences (e.g., `../../`), the rm operation can delete arbitrary files outside the intended task directory. This is reachable from `TaskTracker.delete` which is called by user-facing task management endpoints. Trace: user-supplied task_id -> `_task_path` -> `agfs.rm(force=True)`.
Suggested Fix
Validate all three identifiers against a strict regex and reject any path separators or traversal sequences before constructing the path.
HIGHPath traversal in task store via unvalidated account/user/task IDs
openviking/service/task_store.py:138
[AGENTS: Sentinel]input_validation
`_task_path` concatenates `account_id`, `user_id`, and `task_id` directly into a filesystem path without sanitization. These values are passed from `TaskTracker` calls that originate from user-controlled request parameters (e.g., task_id in API routes). An attacker who can supply a `task_id` like `../../../../etc/cron.d/evil` or a `user_id` with path separators can write or read arbitrary files under the AGFS root. The `_write_task` and `get` methods use this path for `agfs.write` and `agfs.read`. Trace: user-supplied task_id -> `_task_path` -> `agfs.write/read`.
Suggested Fix
Validate account_id, user_id, and task_id against a strict identifier regex (e.g., `^[a-zA-Z0-9_-]{1,128}$`) and reject any path separators, null bytes, or traversal sequences.
HIGHPrivileged RequestContext built from unvalidated queue message fields
openviking/service/user_deletion.py:282
[AGENTS: Sentinel]input_validation
In `_process`, the `target_account_id` and `target_user_id` come directly from the queue message parsed by `_parse_message`. These values are used to construct a `RequestContext` with `Role.ROOT` and then passed to `viking_fs.rm(canonical_user_root(cleanup_ctx), recursive=True)` and `vector_store.delete_user_data`. If an attacker can inject a message with a crafted `target` (e.g., `../../` or a system path), the canonical_user_root could resolve outside the intended user directory, leading to arbitrary file deletion. The fields are not sanitized against path traversal or length limits. Trace: queue message target -> `_process` -> `RequestContext(role=ROOT)` -> `viking_fs.rm`.
Suggested Fix
Validate target_account_id and target_user_id against a strict identifier regex (e.g., `^[a-zA-Z0-9_-]{1,64}$`) and reject any path separators or traversal sequences before constructing the privileged context.
HIGHUnvalidated queue message payload allows arbitrary user deletion
openviking/service/user_deletion.py:478
[AGENTS: Sentinel]input_validation
The `_parse_message` method in `_UserDeletionProcessor` deserializes a queue message without validating the source or authenticating the sender. The `data` dict comes from the queuefs named queue, which is populated by `_deletion_message` but can also be enqueued by any internal component. The message contains `account_id`, `user_id`, and `target` fields that are used directly in `_process` to construct a `RequestContext` with `Role.ROOT` and delete the target user's data. An attacker who can enqueue a message to the USER_DELETION queue (e.g., via an unauthenticated queuefs write path) can delete any user's data. The message fields are only type-checked (str), not validated against the actor's permissions or the deletion fence. Trace: queue message -> `_parse_message` -> `_process` -> `cleanup_ctx = RequestContext(role=Role.ROOT)` -> `viking_fs.rm(canonical_user_root(cleanup_ctx), recursive=True)`.
Suggested Fix
Validate that the message was enqueued by an authorized actor (e.g., include a signed token or verify the queue producer identity), and validate account_id/user_id against a safe character set and length limits before constructing the privileged RequestContext.
HIGHUser deletion task can be processed concurrently, causing data races and duplicate deletion
openviking/service/user_deletion.py:510
[AGENTS: Chaos]Concurrency
The `_UserDeletionProcessor.on_dequeue` schedules `_process` on the service loop via `run_coroutine_threadsafe`. If the queue delivers the same message more than once (e.g., due to a crash after ack, or a retry), or if `initialize()` re-enqueues a task that is already being processed, two concurrent `_process` coroutines can run for the same target user. Both will call `_cancel_user_tasks`, `viking_fs.rm`, and `finish_user_deletion` without any lock or idempotency guard beyond the tracker status check, which is itself racy (read-then-act). This can lead to partial deletion, double-deletion of shared resources, or a crash from concurrent filesystem operations. The `_request_lock` only guards `delete_user`, not `_process`.
Suggested Fix
Add a per-target-user asyncio.Lock (or a set of in-flight task IDs) checked inside `_process` before starting, and re-check the tracker status after acquiring it. Ensure only one `_process` runs per (account_id, user_id) at a time.
HIGHAPI key written into policy set metadata and persisted to disk
openviking/session/train/batch_runner.py:577
[AGENTS: Chaos - Lockdown - Supply - Tripwire]secrets, secrets_management, supply-chain/secret-handling, supply_chain_security
**Perspective 1:** The OpenViking root or user API key is stored in the `_policy_set_metadata` dict under `openviking_api_key`. This metadata is attached to the `ExperienceSet` and later serialized into the run report JSON (`_write_report`) and the train rollout cache files (`_write`), which are written to the repository's `result/` directory. Anyone with read access to those artifacts (e.g., a CI log, shared volume, or backup) obtains a live credential that can impersonate the configured account/user. The key is attacker-reachable only if the artifacts leak, but the exposure is direct and unnecessary. **Perspective 2:** The batch train/eval runner writes the OpenViking API key into the policy set metadata (`_policy_set_metadata`), which is then serialized into the run report and event logs. This exposes the credential in plaintext in artifacts that may be shared, archived, or committed, increasing the risk of credential leakage and unauthorized access to the OpenViking server. The key originates from config or environment and flows into metadata that is persisted via `_write_report` and event recording. **Perspective 3:** `_policy_set_metadata` includes the raw `openviking_api_key` (from `client._api_key`) in the policy set metadata. This metadata flows into `ExperienceSet.metadata`, which is used in `_write_report` (report.json) and `event_recorder.record` (events.jsonl). The API key is therefore persisted to disk in plaintext in the report and event files. Anyone with read access to the result directory (or the events file) can extract the OpenViking API key, which may be a root or user key granting access to the server. This is a credential leak to disk. **Perspective 4:** The OpenViking API key is written into the policy set metadata dict, which is later serialized to disk in the run report and potentially persisted in the experience set. This exposes the credential in plaintext at rest, allowing anyone with filesystem access to the result directory to retrieve it. The key originates from config.server_url/api_key or server_config.root_api_key and is passed to the client, then copied into metadata.
Suggested Fix
Remove the API key from metadata; store only a reference or hash, and ensure credentials are never written to report/event files.
HIGHDataset service endpoints have no authentication
openviking/session/train/components/dataset_service.py:368
[AGENTS: Lockdown - Supply]configuration, supply_chain
**Perspective 1:** The dataset service exposes `/v1/cases/query`, `/v1/rollouts/execute`, and `/v1/rollouts/executions/{execution_id}` without any authentication or authorization. Any unauthenticated caller can query dataset cases, trigger expensive rollouts (consuming compute/LLM tokens), and read rollout results including full messages, tool outputs, prompts, and reasoning. This is an unauthenticated data-exposure and resource-abuse vector. **Perspective 2:** The generic dataset service exposes `/v1/cases/query`, `/v1/rollouts/execute`, and `/v1/rollouts/executions/{id}` without any authentication or authorization. Any network client that can reach the service can submit arbitrary rollouts (consuming compute/LLM tokens) and read rollout results, which may contain sensitive conversation data. There is also no rate limiting, enabling resource-exhaustion attacks. The service is intended to be a remote benchmark host, but the lack of auth means it cannot be safely exposed beyond localhost.
Suggested Fix
Add authentication (e.g., API key or mutual TLS) to all service endpoints, and restrict access to trusted clients.
HIGHRollout execution accepts arbitrary policy set and case from unauthenticated client
openviking/session/train/components/dataset_service.py:388
[AGENTS: Supply]supply_chain
The `/v1/rollouts/execute` endpoint accepts a full `policy_set` (including policy content) and `case` from the client and passes them directly to the rollout executor. An unauthenticated attacker can submit a crafted policy set that causes the executor to load arbitrary content or execute unintended logic, and can consume unbounded compute/LLM resources. The policy content is trusted without integrity verification or provenance checks, and the execution context is also client-controlled.
Suggested Fix
Authenticate clients, validate that policy sets come from a trusted store, and enforce resource limits on rollouts.
HIGHOVPack import reads ZIP without size or member count limits
openviking/storage/ovpack/operations.py:295
[AGENTS: Lockdown - Sentinel - Supply]config_security, input_validation, supply_chain
**Perspective 1:** The `import_ovpack` function opens a user-supplied `.ovpack` file with `zipfile.ZipFile` without enforcing limits on the total archive size, number of members, or individual member sizes. An attacker could provide a maliciously crafted ZIP (e.g., a zip bomb) that decompresses to an enormous size, causing disk exhaustion or memory DoS during extraction and processing. The `validated_import_members` function may validate paths but does not limit sizes. **Perspective 2:** The `import_ovpack` and `restore_ovpack` functions open a ZIP archive from a user-supplied path without verifying a digital signature or a trusted checksum. The archive contents are extracted and written into the Viking filesystem, and vector snapshots are restored. A malicious or tampered .ovpack file could inject arbitrary content into the user's context store. While the manifest includes SHA-256 hashes for individual files, there is no verification of the archive's overall authenticity or provenance. **Perspective 3:** The `import_ovpack` function opens a ZIP file without any size limit on the archive or its members. A malicious or oversized `.ovpack` file could cause excessive memory or disk usage during extraction, leading to denial of service. This is a configuration hardening gap.
Suggested Fix
Enforce limits on the total uncompressed size, number of members, and individual member sizes before extraction. Reject archives that exceed configured thresholds.
HIGHOVPack import reads member data without size limit
openviking/storage/ovpack/operations.py:349
[AGENTS: Chaos - Sentinel]input_validation, resource-exhaustion
**Perspective 1:** In `import_ovpack`, the code reads each ZIP member with `zf.read(safe_zip_path)` without checking the member's uncompressed size. A malicious `.ovpack` file could contain a member that decompresses to an enormous size, causing memory exhaustion when the data is loaded into memory and then written to the filesystem. **Perspective 2:** In `import_ovpack`, the code reads each file member from the ZIP archive entirely into memory with `data = zf.read(safe_zip_path)` before writing it to the filesystem. A malicious or corrupted `.ovpack` file can contain a member that is extremely large (e.g., a multi-GB file), causing the server to allocate an unbounded amount of memory, leading to OOM and a crash. This is a DoS vector if an attacker can upload a crafted `.ovpack` file. The same pattern exists in `restore_ovpack` at line 713.
Suggested Fix
Stream the file member to the destination in chunks (e.g., using `shutil.copyfileobj` with a bounded buffer) and enforce a maximum member size.
HIGHOVPack restore reads ZIP without size or member count limits
openviking/storage/ovpack/operations.py:644
[AGENTS: Sentinel]input_validation
The `restore_ovpack` function opens a user-supplied `.ovpack` file with `zipfile.ZipFile` without enforcing limits on the total archive size, number of members, or individual member sizes. An attacker could provide a zip bomb, causing disk exhaustion or memory DoS during restore.
Suggested Fix
Enforce limits on the total uncompressed size, number of members, and individual member sizes before extraction.
HIGHOVPack restore reads member data without size limit
openviking/storage/ovpack/operations.py:713
[AGENTS: Sentinel]input_validation
In `restore_ovpack`, the code reads each ZIP member with `zf.read(safe_zip_path)` without checking the member's uncompressed size. A malicious backup file could contain a member that decompresses to an enormous size, causing memory exhaustion.
Suggested Fix
Check `member.file_size` before reading and reject members that exceed a configured maximum size.
HIGHOVPack validation reads every file into memory without size limit (ZIP bomb)
openviking/storage/ovpack/validation.py:440
[AGENTS: Chaos]resource_management
validate_manifest_content reads each file member fully into memory via zf.read() to verify size and sha256. A malicious OVPack with a large file (e.g., 10GB) or many files will exhaust memory. There is no per-file or total size cap. This is a ZIP bomb DoS vector when importing an untrusted OVPack.
Suggested Fix
Stream file contents in chunks and enforce a max total size or per-file size limit.
HIGHUnbounded struct format string from untrusted manifest dimensions
openviking/storage/ovpack/vectors.py:155
[AGENTS: Chaos]memory_corruption
In read_dense_vectors, `dimensions` comes from the OVPack manifest (attacker-controlled if the package is untrusted). The format string `f"<{dimensions}f"` is built directly from this value. A malicious manifest can set dimensions to a huge number (e.g. 2^31), causing struct.unpack_from to attempt allocating and reading gigabytes of memory, or raise an unhandled struct.error on truncated data. This can crash the import process or exhaust memory.
Suggested Fix
Validate `dimensions` is a small positive integer (e.g. <= 4096) and that `offset + dimensions*4 <= len(data)` before unpacking.
HIGHQueue message account_id/user_id not validated before constructing privileged RequestContext
openviking/storage/queuefs/add_resource_processor.py:105
[AGENTS: Sentinel]input_validation
The `AddResourceProcessor` dequeues messages from a durable queue and constructs a `RequestContext` directly from `msg.account_id`, `msg.user_id`, `msg.role`, `msg.group_ids`, and `msg.bypass_acl` without validating that these values are trusted or that the message was authenticated. If an attacker can enqueue a message (e.g., via an unauthenticated or weakly-authenticated queue producer, or by replaying/forging a message), they can set `bypass_acl=true` and an arbitrary `role` to escalate privileges and bypass ACL checks during resource processing. The `RequestContext` is then used for filesystem operations (`viking_fs`) and resource service calls, enabling unauthorized read/write/delete of resources.
Suggested Fix
Validate and authenticate queue messages before processing. Ensure `account_id`, `user_id`, `role`, and `bypass_acl` are derived from a trusted authenticated source (e.g., signed message or verified session), and reject messages with `bypass_acl=true` unless the producer is authorized.
HIGHCancelled queue message constructs privileged RequestContext from unvalidated fields
openviking/storage/queuefs/add_resource_processor.py:293
[AGENTS: Sentinel]input_validation
In `on_cancelled`, the code parses the queue message payload and constructs a `RequestContext` with `bypass_acl` and `role` taken directly from the message without validation. An attacker who can inject or forge a cancellation message can set `bypass_acl=true` and an elevated role, causing the processor to release locks or clean up staged sources with elevated privileges. This can lead to unauthorized deletion of staged resources or lock manipulation.
Suggested Fix
Validate and authenticate cancellation messages before constructing the RequestContext. Reject messages with `bypass_acl=true` unless the producer is authorized, and verify role/account/user fields against a trusted source.
HIGHSemantic processor constructs RequestContext with bypass_acl=True from untrusted queue message
openviking/storage/queuefs/semantic_processor.py:167
[AGENTS: Supply]supply_chain
The `_ctx_from_semantic_msg` method builds a `RequestContext` with `bypass_acl=True` directly from fields in a `SemanticMsg` (account_id, user_id, role, group_ids) that are deserialized from the queue payload. If an attacker can inject or tamper with a message in the semantic queue (e.g., via a compromised queue producer or a supply chain attack on a component that enqueues messages), they can set an arbitrary role (e.g., ROOT) and bypass ACL checks for file operations performed during semantic processing. This allows unauthorized access to resources that should be protected by ACLs.
Suggested Fix
Validate the role and identity fields against the authenticated caller's context, and do not set bypass_acl=True for messages originating from untrusted sources.
HIGHSemantic queue message context always bypasses ACL checks
openviking/storage/queuefs/semantic_processor.py:172
[AGENTS: Lockdown]config
_ctx_from_semantic_msg constructs a RequestContext with bypass_acl=True for every message. This means semantic processing operations (file reads, writes, directory listings) skip ACL enforcement entirely. If a queue message is crafted with another user's account_id/user_id, the processor would operate on that user's files without authorization. The message fields are trusted without verification against the actual authenticated identity.
Suggested Fix
Remove bypass_acl=True and verify the message's account/user identity against the authenticated session before processing.
HIGHSemantic queue message deserialized without integrity verification
openviking/storage/queuefs/semantic_processor.py:320
[AGENTS: Supply]supply_chain
The `on_dequeue` method deserializes a `SemanticMsg` directly from the queue payload (`data`) without any integrity verification (e.g., signature, checksum, or provenance). The message contains fields like `account_id`, `user_id`, `role`, `group_ids`, `target_uri`, and `lock_handoff` that are used to construct a `RequestContext` with `bypass_acl=True` and to perform filesystem operations. An attacker who can inject or tamper with a message in the semantic queue can set an arbitrary role (e.g., ROOT) and bypass ACL checks, potentially reading or writing files they should not access. This is a supply chain integrity gap because the queue is trusted without verification.
Suggested Fix
Add integrity verification (e.g., HMAC or signature) to semantic messages before deserialization, and validate the role/identity fields against the authenticated caller.
HIGHCancelled semantic message deserialized without integrity verification and lock released
openviking/storage/queuefs/semantic_processor.py:571
[AGENTS: Lifeline - Supply]error_handling, supply_chain
**Perspective 1:** The `on_cancelled` method deserializes a `SemanticMsg` from the queue payload and, if `msg.lock_handoff` is present, adopts and releases the pathlock lease. An attacker who can inject or tamper with a message in the semantic queue can supply a crafted `lock_handoff` to release a lock they do not own, potentially causing a denial of service or allowing unauthorized access to a locked resource. There is no integrity verification on the message payload. **Perspective 2:** In `on_cancelled`, if the pathlock adopt/release fails (e.g., the lock was already released or the AGFS is unavailable), the exception is caught and only logged as a warning. The method then reports success. This means a lock that was not actually released remains held, potentially blocking subsequent operations on the same path indefinitely. The caller has no way to know the lock was not released.
Suggested Fix
If lock release fails, report an error (not success) so the queue manager can retry the cancellation or escalate. Consider a best-effort retry of the release.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (fetch)
openviking/storage/vectordb/collection/http_collection.py:318
[AGENTS: Supply]supply_chain
The fetch_data method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (delete)
openviking/storage/vectordb/collection/http_collection.py:354
[AGENTS: Supply]supply_chain
**Perspective 1:** The delete_data method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit. **Perspective 2:** The delete_all_data method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (search)
openviking/storage/vectordb/collection/http_collection.py:397
[AGENTS: Supply]supply_chain
The search_by_vector method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (search by id)
openviking/storage/vectordb/collection/http_collection.py:440
[AGENTS: Supply]supply_chain
The search_by_id method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (multi-modal search)
openviking/storage/vectordb/collection/http_collection.py:484
[AGENTS: Supply]supply_chain
The search_by_multimodal method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (random search)
openviking/storage/vectordb/collection/http_collection.py:527
[AGENTS: Supply]supply_chain
The search_by_random method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (keyword search)
openviking/storage/vectordb/collection/http_collection.py:570
[AGENTS: Supply]supply_chain
The search_by_keywords method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (scalar search)
openviking/storage/vectordb/collection/http_collection.py:616
[AGENTS: Supply]supply_chain
The search_by_scalar method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHVikingDB HTTP collection client uses plaintext HTTP without TLS (aggregate)
openviking/storage/vectordb/collection/http_collection.py:658
[AGENTS: Supply]supply_chain
The aggregate_data method constructs its API endpoint using plaintext HTTP. This means all data transmitted to and from the vector database — including vector embeddings, search queries, and potentially sensitive metadata — is sent in cleartext. An attacker on the network path can intercept, modify, or inject data, compromising the integrity and confidentiality of the vector store. This is a supply chain integrity issue because the client cannot verify the authenticity of the server or the integrity of the data in transit.
Suggested Fix
Use HTTPS for all VikingDB API endpoints and enforce TLS certificate verification.
HIGHBytesRow.serialize builds unbounded struct format strings from untrusted list lengths
openviking/storage/vectordb/engine/_python_api.py:217
[AGENTS: Chaos]resource_exhaustion
OWASP A04:2021NIST SC-5
The `serialize` method constructs a `struct` format string by interpolating the length of user-supplied list fields (e.g. `f"{len(items)}q"`). A caller can pass a list with millions of elements, causing `struct.calcsize(fmt)` to allocate a huge buffer and the format string to grow without bound, exhausting memory. The input originates from data-plane API requests (e.g. `upsert_data`), so an authenticated user can trigger OOM. There is no cap on list length or total serialized size.
Suggested Fix
Validate and cap the number of items per list field and the total serialized row size before building the format string.
HIGHDrop collection endpoint has no authorization check
openviking/storage/vectordb/service/api_fastapi.py:218
[AGENTS: Chaos - Sentinel]authorization, input_validation
**Perspective 1:** The `/DeleteVikingdbCollection` endpoint calls `project.drop_collection()` with no authentication or authorization. Any caller can delete any collection by name. This is a data-loss vulnerability. The endpoint is registered on the collection_router with no auth dependency. **Perspective 2:** The DeleteVikingdbCollection endpoint accepts a `CollectionName` from the HTTP request and passes it directly to `project.drop_collection` without validating the name. An attacker can supply a collection name with path traversal sequences or special characters, potentially causing the backend to delete files or directories outside the intended collection storage. The endpoint also lacks an authorization check to ensure the caller is allowed to drop the collection.
Suggested Fix
Validate the `CollectionName` against a safe character set and enforce authorization before dropping the collection.
HIGHUnvalidated upsert data allows arbitrary field injection and type confusion
openviking/storage/vectordb/service/api_fastapi.py:235
[AGENTS: Sentinel]input_validation
The /api/vikingdb/data/upsert endpoint accepts a `fields` payload from the HTTP request and passes it directly to `collection.upsert_data` without validating the data against the collection's schema. An attacker can send arbitrary field names, types, and values (including oversized strings, nested objects, or vectors of unexpected dimensions), potentially corrupting stored records, bypassing schema constraints, or causing the backend to misinterpret data types. The `data_utils.convert_dict` only performs a shallow conversion and does not enforce the collection's field definitions. This is a direct input-validation gap at the API boundary.
Suggested Fix
Validate the incoming `fields` payload against the collection's schema (field names, types, dimensions, and value ranges) before calling `upsert_data`. Reject unknown fields and enforce type/length constraints.
HIGHDelete all data endpoint has no authorization check
openviking/storage/vectordb/service/api_fastapi.py:280
[AGENTS: Chaos]authorization
The `/api/vikingdb/data/delete` endpoint with `del_all=true` calls `collection.delete_all_data()`. There is no authentication or authorization check on this endpoint — any caller who can reach the API can wipe an entire collection. This is a data-loss vulnerability. The endpoint is registered on the data_router with no auth dependency.
Suggested Fix
Add authentication and authorization checks to all data endpoints, especially delete_all.
HIGHUnvalidated string length in deserialize_field allows out-of-bounds read
openviking/storage/vectordb/store/bytes_row.py:253
[AGENTS: Chaos - Sentinel]data_integrity, input_validation
**Perspective 1:** The deserialize_field method reads a string length (str_len) from the serialized data without validating that the offset and length are within the bounds of the serialized_data buffer. An attacker who can control the serialized row data (e.g., via a crafted database file or API payload) can supply a large length value, causing the subsequent slice serialized_data[str_offset : str_offset + str_len] to read beyond the buffer. This can lead to out-of-bounds memory access, information disclosure, or a crash. The same pattern is repeated for binary, text, and list_string fields. **Perspective 2:** In deserialize_field, for string/binary/text/list types, the offset and length values are read directly from the serialized_data buffer without validating that they are within the buffer bounds. A corrupted or maliciously crafted row (e.g., from a partially written file or a tampered store) can supply an offset/length that causes struct.unpack_from to raise struct.error or return data from an unintended region, or cause an IndexError. This can crash the deserialization path with an uncaught exception, breaking store reads. The input path is any stored row that is read back from disk, which can be corrupted by a crash mid-write or by an attacker with write access to the store.
Suggested Fix
Before unpacking, validate that str_offset + str_len <= len(serialized_data) and that offsets are >= 0. Add explicit bounds checks for all variable-length field reads.
HIGHUnvalidated list length in deserialize_field allows out-of-bounds read
openviking/storage/vectordb/store/bytes_row.py:268
[AGENTS: Sentinel]input_validation
The deserialize_field method reads a list length (list_len) from the serialized data without validating that the list_offset and the total size of the list elements are within the bounds of the serialized_data buffer. An attacker controlling the serialized data can set a large list_len, causing the loop to read beyond the buffer (for list_string) or struct.unpack_from to read out of bounds (for list_int64 and list_float32). This can lead to memory disclosure or a crash.
Suggested Fix
Validate that list_offset + (list_len * element_size) <= len(serialized_data) before iterating or unpacking.
HIGHReDoS via user-supplied regex pattern in grep
openviking/storage/viking_fs/_grep.py:326
[AGENTS: Chaos - Sentinel]denial_of_service, input_validation
**Perspective 1:** The `pattern` parameter is user-controlled (from the HTTP API / CLI) and is compiled directly with `re.compile` without any length or complexity limits. A malicious pattern with nested quantifiers (e.g. `(a+)+$`) can cause catastrophic backtracking, leading to CPU exhaustion (ReDoS). This affects the `_grep_in_files` path used by the vikingdb_then_fs engine. The same issue exists at line 485 in `_grep_encrypted`. **Perspective 2:** The `pattern` parameter is user-controlled (from the `grep` API) and is compiled directly with `re.compile` without any complexity limits. A malicious pattern like `(a+)+$` against a long string of 'a's can cause catastrophic backtracking, consuming 100% CPU and blocking the event loop. This is reachable via the `_grep_in_files` and `_grep_encrypted` paths, which process every line of every candidate file.
Suggested Fix
Validate the pattern against a maximum length and reject patterns containing nested quantifiers, or use a regex engine with linear-time guarantees (e.g. `re2`).
HIGHReDoS via user-supplied regex pattern in encrypted grep
openviking/storage/viking_fs/_grep.py:485
[AGENTS: Sentinel]input_validation
The `pattern` parameter is user-controlled and compiled directly with `re.compile` in the `_grep_encrypted` path. A malicious pattern with catastrophic backtracking (e.g. `(a|aa)+$`) can cause CPU exhaustion. This path is reached when encryption is enabled or when agfs native grep is unavailable.
Suggested Fix
Apply the same regex validation as the fs path: enforce a max length and reject nested-quantifier patterns, or use a linear-time regex engine.
HIGHdelete_account_data can delete any account's data via proxy
openviking/storage/vikingdb_manager.py:498
[AGENTS: Chaos]authorization
`VikingDBManagerProxy.delete_account_data` passes `account_id` directly to the manager with the proxy's bound ctx. If the caller can control `account_id` (e.g. from a request), they can delete another tenant's data. The proxy binds a ctx but the method takes an explicit account_id parameter, which may not be validated against the ctx's account. This is a cross-tenant data deletion vulnerability.
Suggested Fix
Validate account_id matches the bound ctx account, or remove the parameter.
HIGHInstall script piped from remote URL to bash without integrity verification
openviking_cli/utils/ollama.py:146
[AGENTS: Lockdown - Supply - Tripwire]config, supply_chain
**Perspective 1:** The install_ollama function downloads and executes a shell script from https://ollama.com/install.sh by piping it directly to `sh` via `bash -c`. The script is fetched over HTTPS but there is no integrity check (checksum, signature, or pinned version). If the remote endpoint is compromised or serves a malicious script, arbitrary code is executed with the privileges of the current user. This is a supply chain risk that can lead to RCE during installation. The same pattern is used for both Darwin and Linux platforms. **Perspective 2:** The install_ollama function pipes a remote install script directly into a shell. The script is fetched over HTTPS but its content is not pinned to a known hash, signature, or version. A compromised or tampered script at ollama.com (or a MITM if TLS is bypassed) would execute arbitrary code with the privileges of the OpenViking server process. This is a classic remote-code-execution supply chain vector: untrusted network content -> subprocess shell execution. **Perspective 3:** The `install_ollama` function executes `curl -fsSL https://ollama.com/install.sh | sh` via a shell. This downloads and executes a remote script without verifying its integrity (checksum/signature). If the download is intercepted or the upstream is compromised, arbitrary code is executed with the privileges of the user running the CLI.
Suggested Fix
Download the installer to a temporary file, verify its SHA-256 checksum against a known-good value, and only then execute it. Alternatively, use a package manager (e.g., brew) with pinned versions.
HIGHOllama install script downloaded and executed without integrity verification (Linux)
openviking_cli/utils/ollama.py:153
[AGENTS: Supply]supply_chain
Same as the macOS path: the Linux install path also pipes an unverified remote script into a shell. The script is fetched over HTTPS but its content is not pinned to a known hash, signature, or version. A compromised or tampered script at ollama.com (or a MITM if TLS is bypassed) would execute arbitrary code with the privileges of the OpenViking server process. This is a classic remote-code-execution supply chain vector: untrusted network content -> subprocess shell execution.
Suggested Fix
Pin the installer to a specific version and verify its SHA-256 checksum before execution, or distribute the installer as a signed package via the OS package manager.
MEDIUMLDAP filter injection via user_search_filter template
openviking/server/auth/plugins/ldap.py:223
[AGENTS: Sentinel]input_validation
The `user_search_filter` is a configurable template that is formatted with the username. While `ldap.filter.escape_filter_chars` is applied to the username, the template itself is not validated. If an administrator configures a filter that does not use the `%s` placeholder (e.g., a static filter or one that concatenates the username in an unsafe way), an attacker-controlled username could inject LDAP filter syntax. More critically, the template is applied with `%` formatting, which can raise a `TypeError` or `ValueError` if the username contains `%` characters, potentially causing a denial of service. The input path is: HTTP request -> `_extract_credentials` -> `_authenticate_ldap` -> `_search_user_attributes` -> this line.
Suggested Fix
Validate that `user_search_filter` contains exactly one `%s` placeholder and no other format specifiers. Also, ensure the username is always passed through `ldap.filter.escape_filter_chars` and that the template is a static string from config, not user-controlled.
MEDIUMAuthentication falls back to unauthenticated dev mode by default
openviking/server/config.py:335
[AGENTS: Lockdown]Insecure Configuration
When `auth_mode` is not set and `root_api_key` is not configured, `get_effective_auth_mode()` returns `AuthMode.DEV.value`, which grants unauthenticated ROOT access (as reported in the dev auth plugin). This means a default deployment with no explicit `root_api_key` exposes the entire API without any authentication, allowing anyone to read, modify, or delete all data. This is a critical insecure default that should require explicit configuration of an auth mode or root key.
Suggested Fix
Require `root_api_key` or an explicit `auth_mode` at startup, and refuse to start in dev mode for production deployments.
MEDIUMPermissive CORS allows any origin to access the API
openviking/server/config.py:341
[AGENTS: Lockdown]Insecure Configuration
**Perspective 1:** The server config defaults `cors_origins` to `["*"]`, which allows any website to make authenticated cross-origin requests to the OpenViking API. Since the API uses API keys sent via the `X-API-Key` header, a malicious site could trigger state-changing requests (e.g., resource deletion, session commits) on behalf of a logged-in user if the user visits the malicious page. This is a direct CORS misconfiguration that enables cross-site request forgery against the API. **Perspective 2:** The default `cors_origins` of `["*"]` combined with the API's use of `X-API-Key` headers means any malicious website can make authenticated requests to the OpenViking API. Since the API key is stored in sessionStorage by the web studio, a cross-origin request from a malicious page could read or modify user data if the user is logged in. This is a direct CORS misconfiguration enabling cross-site request forgery and data theft.
Suggested Fix
Default `cors_origins` to an empty list or a specific allowlist of trusted origins, and require explicit configuration to enable cross-origin access.
MEDIUMAPI key hashing disabled by default stores keys in plaintext
openviking/server/config.py:345
[AGENTS: Lockdown]Insecure Configuration
The server config defaults `api_key_hashing_enabled` to `False`. When file-level encryption is disabled (the default for `encryption_enabled`), API keys are stored in plaintext on disk. Even when file encryption is enabled, the code explicitly logs that keys are stored in plaintext within AES-GCM encrypted files. An attacker with filesystem access (e.g., via a compromised container or backup) can read all API keys and impersonate any user. This is an insecure default that should enable hashing by default.
Suggested Fix
Default `api_key_hashing_enabled` to `True`, or require explicit opt-in to store keys in plaintext.
MEDIUMRole downgrade gate silently bypassed when role resolver raises
openviking/server/oauth/provider.py:300
[AGENTS: Chaos - Lifeline - Supply - Tripwire]auth, error_handling, supply_chain
**Perspective 1:** In exchange_refresh_token, the role downgrade check wraps the role resolver call in a broad except (ValueError, Exception) that sets both token_role and current_role to None. If the role resolver raises (e.g., transient DB error, malformed stored role), the check is skipped entirely and the refresh proceeds. This defeats the security invariant that a demoted user cannot mint fresh access tokens at their old privilege level. An attacker who previously held a high-privilege refresh token can keep rotating it indefinitely whenever the resolver errors, or can trigger the error condition (e.g., by corrupting the stored role string) to bypass the gate. The except clause should only catch ValueError from Role() parsing, not swallow resolver failures; a resolver failure should fail closed (reject the refresh). **Perspective 2:** The role downgrade gate in exchange_refresh_token catches (ValueError, Exception) and sets both token_role and current_role to None. This means any exception during role resolution (e.g. a transient DB error from the role_resolver) silently disables the privilege-escalation protection, allowing a demoted user to continue minting access tokens at their old higher role. The catch-all also swallows programming errors. The input is the refresh_token.role and the role_resolver result; the sink is the subsequent rank comparison that is skipped when both are None. This is a security-relevant error handling flaw because a failure to resolve the current role should fail closed (reject the refresh), not fail open. **Perspective 3:** In exchange_refresh_token, the role downgrade gate catches (ValueError, Exception) and sets both roles to None, which then skips the rank comparison. This means any exception (including a transient DB error from the role_resolver) silently disables the privilege-escalation protection, allowing a demoted user to continue minting access tokens at their old higher role. The input is the refresh_token.role and the role_resolver result; the sink is the skipped rank check. **Perspective 4:** In `exchange_refresh_token`, the `role_resolver` callback is wrapped in a broad `except (ValueError, Exception)` that sets both `token_role` and `current_role` to `None`. This means that if the role resolver fails (e.g., due to a database error or a misconfigured user), the role downgrade gate is silently bypassed, and the refresh token is rotated without enforcing the role check. An attacker who can cause the role resolver to fail (e.g., by triggering an exception) could potentially obtain a refresh token at a higher privilege level than their current role. **Perspective 5:** The role resolution in exchange_refresh_token catches all exceptions and sets both roles to None. This could mask errors and potentially allow a role downgrade check to be bypassed if the role resolver fails. This is a defense-in-depth gap.
Suggested Fix
Catch only ValueError for invalid role strings, and let other exceptions propagate so the refresh fails closed. If a transient resolver error must be tolerated, treat it as a downgrade (reject) rather than as a pass.
MEDIUMMissing authorizing_key_fp stored as empty string, not NULL, weakening key-rotation binding
openviking/server/oauth/provider.py:405
[AGENTS: Chaos - Lifeline - Lockdown - Supply]auth, configuration_security, error_handling, supply_chain
**Perspective 1:** The comment claims a missing fp makes the token unusable because the bearer-auth path fails closed against a NULL stored value. But this line converts None to an empty string before insert_access/insert_refresh. If the storage layer stores '' rather than NULL, the bearer-auth check that compares against a NULL may not trigger, and the token could be accepted without a valid authorizing key fingerprint. This weakens the security property that rotating the API key invalidates the entire OAuth token chain. An attacker who obtains a token minted through a path that omitted the fp (e.g., a legacy or malformed flow) could keep using it even after the authorizing key is rotated. **Perspective 2:** In _issue_token_pair(), if authorizing_key_fp is None (e.g., a legacy flow or a code path that doesn't record it), the code substitutes an empty string and still issues a valid access/refresh token pair. The comment acknowledges that a missing fp makes the token 'unusable' in the bearer-auth path, but the token is still stored and returned to the client. If the bearer-auth path fails open (or if another consumer trusts the token without checking fp), an attacker could obtain a token that is not bound to any API key, defeating key-rotation revocation. This is a fail-open configuration for a security-critical binding. **Perspective 3:** In _issue_token_pair, when authorizing_key_fp is None it is coerced to an empty string and stored. The comment acknowledges the bearer-auth path will fail-closed against a NULL stored value, but here an empty string is stored instead of NULL, which may bypass that fail-closed check if the bearer-auth path treats empty string as 'no binding'. This can occur if a code path calls _issue_token_pair without a fingerprint (e.g. a future flow or a bug in the OTP verification). The result is an access token that may authenticate without the API-key binding, weakening the rotation-invalidation guarantee. The input is the authorizing_key_fp parameter; the sink is the stored token record. **Perspective 4:** In `_issue_token_pair`, the `authorizing_key_fp` is set to an empty string (`""`) when it is `None`. This means that if a token is issued without a valid authorizing key fingerprint (e.g., due to a bug or a compromised flow), the token will be stored with an empty fingerprint. The comment notes that the bearer-auth path will fail-closed against a NULL stored value, but an empty string may not be treated as NULL, potentially allowing a token to be used without a valid authorizing key binding. This could weaken the security of the token chain.
Suggested Fix
Raise a clear error when authorizing_key_fp is None instead of silently substituting an empty string, or ensure the bearer-auth path treats empty string identically to NULL and rejects it.
MEDIUMSSRF — user-controlled URL in HTTP request
openviking/server/oauth/router.py:133
[AGENTS: rules-engine]security
HTTP request with user-controlled URL in openviking/server/oauth/router.py at line 133 enables SSRF attacks.
Suggested Fix
Validate URLs against an allowlist. Block private/internal IP ranges.
MEDIUMGit repo URL passed to subprocess without shell injection guard
openviking/server/openviking_assets.py:329
[AGENTS: Sentinel]input_validation
The repo_url is user-supplied via the manifest/catalog YAML and passed directly to `asyncio.create_subprocess_exec` as an argument. While `create_subprocess_exec` avoids shell interpretation, the URL is not validated against a scheme allowlist beyond `require_remote_resource_source`. A URL like `https://host/../../etc/passwd` or one with a crafted path could cause git to read arbitrary local files or interact with unexpected endpoints. The `_validate_clone_url` only checks for control chars, leading '-', and remote-helper prefixes; it does not restrict the scheme or host. An attacker controlling the manifest can point git at internal network resources (SSRF) or local file paths.
Suggested Fix
Enforce an explicit allowlist of schemes (https, ssh, git) and validate the host against an SSRF blocklist (reject localhost, private IP ranges, metadata endpoints) before invoking git.
MEDIUMAccount ID path traversal in delete_account allows deleting arbitrary AGFS paths
openviking/server/routers/admin.py:404
[AGENTS: Chaos - Lockdown]Data Integrity, Error Handling, Path Traversal / Injection, Race Condition, Race Condition / TOCTOU, Resource Exhaustion, configuration
**Perspective 1:** The `account_id` path parameter is interpolated directly into an AGFS filesystem path without validation or sanitization. An attacker who can reach this endpoint (requires root auth) could supply an account_id containing path traversal sequences such as `../../` or absolute path components. Since `rm` is called with `recursive=True`, this could delete arbitrary directories under the AGFS root, including system or other tenants' data. The input originates from the HTTP path parameter `account_id` and reaches the `rm` sink without any allowlist or normalization. **Perspective 2:** The `delete_account` endpoint constructs an AGFS path directly from the `account_id` path parameter and calls `rm` recursively. If `account_id` contains path traversal characters (e.g., `../`), it could delete data outside the intended account directory. The account_id should be validated against a safe pattern. **Perspective 3:** The `delete_account` endpoint does not acquire any locks on the account being deleted. If a concurrent request is writing to the account (e.g., adding a resource, creating a session) while the deletion is in progress, the deletion could remove data that is being written, or the write could fail with a 'not found' error. This can lead to data loss or inconsistent state. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 4:** The `delete_account` endpoint wraps the AGFS and VectorDB cleanup in broad `try/except` blocks that log a warning and continue. This means if the cleanup fails (e.g., permission denied, disk full, network error), the failure is silently ignored and the account metadata is still deleted, leaving orphaned data. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 5:** The `delete_account` endpoint checks for account existence (via `_check_account_exists`) and then performs the deletion. Between the check and the deletion, another request could create the same account, leading to the deletion of the newly created account's data. This is a time-of-check-to-time-of-use (TOCTOU) race. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 6:** The `delete_account` endpoint does not prevent deletion of the default account or system accounts. If an attacker with root access deletes the default account, it could break the entire system, as many operations may depend on the default account existing. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 7:** The `delete_account` endpoint does not prevent concurrent deletion of the same account. If two requests delete the same account simultaneously, the second request may fail with a 'not found' error, or the two deletions may interfere with each other, causing data corruption. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 8:** The `delete_account` endpoint calls `viking_fs._async_agfs.rm` without a timeout. If the AGFS rm operation blocks (e.g., due to a network issue or a very large directory), the request could hang indefinitely, tying up a worker thread and potentially exhausting the thread pool. The input is the `account_id` path parameter, reaching the `rm` sink. **Perspective 9:** The `delete_account` endpoint does not call `_check_account_exists` before attempting the deletion. If the account does not exist, the AGFS rm and VectorDB delete will likely fail, and the `manager.delete_account` may also fail. The error handling is inconsistent, and the caller may receive a generic 500 error instead of a clear 'not found' message. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 10:** The `delete_account` endpoint deletes the account's data without checking for active sessions or resources. If the account has active sessions, deleting the account could cause those sessions to fail or leave orphaned state. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 11:** The `delete_account` endpoint allows a root caller to delete their own account. This could lock the caller out of the system and cause data loss. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 12:** The `delete_account` endpoint calls `storage.delete_account_data(account_id)` which may need to delete a very large number of vector records. This operation could take a long time or exhaust memory. The input is the `account_id` path parameter, reaching the `delete_account_data` sink. **Perspective 13:** The `delete_account` endpoint calls `viking_fs._async_agfs.rm` with `recursive=True`. If the account has a very large number of files, this operation could take a long time or exhaust memory. The input is the `account_id` path parameter, reaching the `rm` sink. **Perspective 14:** The `delete_account` endpoint deletes the account's data but does not explicitly clean up groups or group members. If the account has a large number of groups, this could leave orphaned group data. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks. **Perspective 15:** The `delete_account` endpoint deletes the account's data but does not explicitly clean up all users in the account. If the account has a large number of users, this could leave orphaned user data. The input is the `account_id` path parameter, reaching the `rm` and `delete_account_data` sinks.
Suggested Fix
Validate account_id against a strict pattern (e.g., `^[a-zA-Z0-9_-]+$`) before using it in filesystem paths, and reject any value containing `/`, `\`, `..`, or null bytes.
MEDIUMOpenViking assets preflight accepts arbitrary git repo_url and credentials without integrity verification
openviking/server/routers/openviking_assets.py:71
[AGENTS: Supply]supply_chain
The `/preflight` endpoint accepts a `repo_url`, `branch`, `commit`, and optional `auth_config` (username/token) from the caller and uses them to verify source readability by connecting to the git repository. An attacker who can call this endpoint can supply an arbitrary `repo_url` (e.g., an attacker-controlled server) and credentials, causing the server to make outbound connections to that server (SSRF) and potentially exfiltrate the provided credentials or probe internal network resources. There is no allowlist of trusted repositories or verification of the repo's integrity.
Suggested Fix
Restrict the `repo_url` to a trusted allowlist of repositories, and do not accept arbitrary credentials from clients for preflight checks.
Note: Fixing issues can create a domino effect — resolving one finding often surfaces new ones that were previously hidden. Multiple scan-and-fix cycles may be needed until you’re satisfied no further issues remain. How deep you go is your call.