refactor: no comments left - one-line module docstrings, contracts on public fields only; jobs/job.py; example config and README

This commit is contained in:
hh
2026-09-02 00:33:04 +02:00
parent 3915ab48f9
commit 0e07e8409d
65 changed files with 795 additions and 1891 deletions
+3 -1
View File
@@ -6,9 +6,11 @@
- **NEVER write comments** - code should be self-documenting - **NEVER write comments** - code should be self-documenting
## Layout ## Layout
- `src/beaver_gateway` - python, unopinionated primitives (no vault paths, no agent names) - `src/beaver_gateway` - python, unopinionated primitives (no vault paths, no agent names, no Russian): one package per capability - `agents`, `backends`, `conversations`, `jobs`, `frontends`, `mcp`, `vault`, `security`, `events`, `storage`; `config.py` loads a setup, `app.py` builds the runtime, `cli.py` is the entrypoint
- every string shown to a model or a user lives in a `*Texts` dataclass with English defaults (`conversations/texts.py`, `frontends/telegram/texts.py`); a setup overrides them
- `ui` - SvelteKit SPA (admin + Obsidian panel), served by the gateway from `ui/build` - `ui` - SvelteKit SPA (admin + Obsidian panel), served by the gateway from `ui/build`
- `docs/MISSION.md`, `docs/PRD.md`, `docs/PLAN.md`, `docs/PROGRESS.md` - keep PROGRESS.md current - `docs/MISSION.md`, `docs/PRD.md`, `docs/PLAN.md`, `docs/PROGRESS.md` - keep PROGRESS.md current
- docstrings: one line per module saying what it is for, one line per public configuration field; no section references, phases or history
## Checking commands ## Checking commands
After writing code, always run (or `make check`): After writing code, always run (or `make check`):
+60 -2
View File
@@ -1,2 +1,60 @@
# Beaver Gateway # beaver-gateway
[![AI Slop Inside](https://sladge.net/badge.svg)](https://sladge.net)
A gateway for personal agents built on the Claude Agent SDK: conversations
with a lifecycle, windows to talk through, jobs, memory hooks, and a policy
boundary - without an opinion about what the agent is for. A setup is one
`config.py` (or a package next to it) that assembles a `Gateway`; this repo
holds the primitives, [beaver-agent](https://git.kotikot.com/beaver/beaver-agent)
is a full setup built on them.
## What it gives a setup
- **agents/** - what an agent is: a Claude agent with a prompt per
conversation kind, skill sets per kind, gateway tools, a `PreToolUse`
policy; a Raycast agent for cheap fast answers.
- **backends/** - how a turn runs: the Claude Agent SDK subprocess with a
session pool, transcripts and session stores; the Raycast wire.
- **conversations/** - the model of a conversation: kinds (master, branch,
deep, job, fork), rows and windows, the inject queue with priorities,
seeds, turns, questions, closing through a distiller, master rotation,
the envelope, the in-process gateway tools.
- **jobs/** - cron, webhook and event jobs on pgqueuer, deferred injects,
the subscription budget.
- **frontends/** - the windows: Telegram (master = General, topic = branch),
markdown files in a vault, `/api` + the admin SPA, an Anthropic-compatible
`/anthropic/v1/messages`, MCP re-exposure at `/mcp/<name>`.
- **mcp/** - MCP servers a setup declares (stdio, http, python tools),
aggregated in-process and handed to agents by name.
- **vault/** - watching a directory of notes for the envelope.
- **security/** - bearer tokens, the audit log, credential redaction.
- **events/** - the in-process bus every frontend taps, and the stream
event protocol backends emit.
- **storage/** - SQLModel tables and the Postgres session store.
`config.py` defines `Gateway` and loads a setup; `app.py` builds the runtime
from it; `cli.py` is the entrypoint. Every string the gateway puts in front of
a model or a user has an English default and is overridable
(`ConversationTexts`, `TelegramTexts`).
## Run
```sh
uv sync # python
cd ui && bun install && cd .. # admin + panel SPA, built into ui/build
cp examples/config.py /config/config.py
DATABASE_URL=postgresql+psycopg://... CONFIG_PATH=/config/config.py uv run beaver-gateway
```
`examples/config.py` is the smallest setup that starts. The setup's own
secrets come from `.env` (`load_dotenv` runs before the config is executed).
## Develop
```sh
make check # ruff format --check, ruff check, ty, pytest
make fix
```
No comments in code: modules carry a one-line docstring saying what they are
for, public configuration fields carry a one-line docstring saying what they
mean; everything else is the code.
+35 -175
View File
@@ -1,204 +1,64 @@
# Sample user config — grows alongside the implementation phases. """The smallest setup: one agent, one tool, one job, the built-in frontends.
#
# Loader (beaver_gateway/config_loader.py) execs this file with The loader execs this file with ``ClaudeAgent``, ``RaycastAgent``,
# ClaudeAgent, RaycastAgent, McpServer, ExposedMcp, Gateway already ``McpServer``, ``ExposedMcp`` and ``Gateway`` already bound; the imports
# bound — so importing them is optional. We import explicitly here so below are for type checkers. A real setup grows into a package next to
# IDEs and type-checkers see real symbols instead of free variables. this file (see beaver-agent).
"""
import logging
import tempfile import tempfile
from datetime import date from datetime import UTC, datetime
from pathlib import Path from pathlib import Path
from beaver_gateway.agents.base import ExposedMcp from beaver_gateway.agents.base import ExposedMcp
from beaver_gateway.agents.claude import ClaudeAgent from beaver_gateway.agents.claude import ClaudeAgent, Prompts
from beaver_gateway.agents.raycast import RaycastAgent, RemoteTool, UserPreferences from beaver_gateway.config import Gateway
from beaver_gateway.app import Gateway
from beaver_gateway.frontends.turn_record import slugify
from beaver_gateway.frontends.admin import AdminFrontend from beaver_gateway.frontends.admin import AdminFrontend
from beaver_gateway.frontends.anthropic import AnthropicMessagesFrontend from beaver_gateway.frontends.anthropic import AnthropicMessagesFrontend
from beaver_gateway.frontends.api import ApiFrontend
from beaver_gateway.frontends.markdown import MarkdownFrontend from beaver_gateway.frontends.markdown import MarkdownFrontend
from beaver_gateway.frontends.mcp_server import McpServerFrontend from beaver_gateway.frontends.mcp_server import McpServerFrontend
from beaver_gateway.jobs.scheduler import Job, JobRun
from beaver_gateway.mcp.types import McpServer from beaver_gateway.mcp.types import McpServer
log = logging.getLogger("example")
def chat_path(title: str, agent: str, vault: Path) -> Path: # noqa: ARG001 VAULT = Path(tempfile.mkdtemp(prefix="beaver-vault-")).resolve()
"""Where a new chat file lands in the vault. PROMPT = VAULT / "assistant.md"
PROMPT.write_text(
Called by ``MarkdownFrontend`` for every conversation that needs a "You are a concise assistant. When asked the time, call `current_time`.\n"
file: a ``deep`` chat spawned by the dispatcher, a ``/v1/messages`` )
chat (``title`` = first user message), or with ``log_all_chats=True``
the archive of a stateless agent's turns. Return value can be
absolute or relative; relative paths are anchored under ``vault``.
Layout below: ``<vault>/<YYYY-MM>/<YYYY-MM-DD>_<topic>.md``.
"""
today = date.today()
return vault / f"{today:%Y-%m}" / f"{today:%Y-%m-%d}_{slugify(title, maxlen=40)}.md"
def current_time() -> str: def current_time() -> str:
"""Return the current local time as an ISO-8601 string. return datetime.now(UTC).astimezone().isoformat()
Trivial demo tool for the Phase 2.1 internal MCP aggregator —
confirms a ``python_tool`` namespace is reachable on
``http://127.0.0.1:<INTERNAL_MCP_PORT>/mcp/time``.
"""
from datetime import datetime
return datetime.now().astimezone().isoformat() async def heartbeat(run: JobRun) -> None:
master = await run.master()
log.info("heartbeat: master=%s", master.external_id if master else None)
gateway = Gateway( gateway = Gateway(
agents=[ agents=[
# Phase 2.2 — ClaudeCodeBackendAdapter routes this agent's
# ``/v1/messages`` calls through ``claude-code-api``. The
# ``time`` MCP gets exposed as ``mcp__time__current_time`` to
# the subscription claude session via
# ``BackendOptions.mcp_servers`` pointing at the internal
# aggregator on ``127.0.0.1:INTERNAL_MCP_PORT/mcp/time/``.
#
# Fresh empty tempdir (not a hardcoded ``/tmp``) for two
# reasons: claude-code-api derives the JSONL project-key from
# ``cwd``, but claude itself writes the JSONL using the cwd's
# realpath — on macOS ``/tmp`` and ``/var/folders/...`` are
# both ``/private/...`` symlinks, so unresolved cwds make
# ``JsonlWatcher`` time out waiting on the wrong path. The
# explicit ``.resolve()`` collapses the symlink before claude
# ever sees the dir, and ``mkdtemp`` guarantees the directory
# is empty so claude does not pick up leftover files.
ClaudeAgent( ClaudeAgent(
name="stub", name="assistant",
model="claude-sonnet-4-6", model="claude-sonnet-5",
# ``system_prompt`` is appended to claude's built-in agent cwd=VAULT,
# prompt (via ``--append-system-prompt``) — so it adds the prompts=Prompts(master=(PROMPT,), branch=(PROMPT,), deep=(PROMPT,)),
# agent's identity on top of claude-code's baseline tool gateway_tools=("spawn", "say", "schedule"),
# knowledge, rather than replacing it. Same shape as the
# RaycastAgent's ``system_prompt → additional_system_instructions``
# mapping. For full ``BackendOptions`` knobs (timeouts,
# extra_args, history mode, etc.) import ``ClaudeCodeOptions``
# and pass ``options=ClaudeCodeOptions(...)``.
system_prompt=(
"You are a stub agent used to validate the Phase 0 skeleton.\n"
"If asked the current time, call the `current_time`"
" MCP tool instead of guessing."
),
cwd=Path(tempfile.mkdtemp(prefix="beaver-stub-cwd-")).resolve(),
expose_mcps=(ExposedMcp(name="time"),), expose_mcps=(ExposedMcp(name="time"),),
), )
# Phase 1.2 — a RaycastAgent the AnthropicMessagesFrontend will
# route via RaycastBackend. Phase 1.5 added the per-agent knobs
# (`temperature` / `additional_system_instructions` / etc.) —
# only `model`, `system_prompt`, and at least one of the others
# is mandatory.
RaycastAgent(
name="research",
model="Gemini 3.1 Flash Lite",
system_prompt=(
"You are a research assistant. "
"Reply in the user's language. Cite URLs when you use web search."
),
temperature=0.5,
available_native_tools=(RemoteTool.WEB_SEARCH, RemoteTool.READ_PAGE),
# Lambda so today's date is rebuilt on every request while
# locale/timezone stay pinned. ``True`` would give the same
# fresh-date behaviour but would also auto-pick host locale
# and timezone (``en-US`` / system tz), which isn't what we
# want here.
user_preferences=lambda: UserPreferences(
locale="en-GB",
timezone="Europe/Berlin",
current_date=date.today().isoformat(), # noqa: DTZ011 — local date is intended
),
),
], ],
mcps=[ mcps=[McpServer.python_tool(name="time", tools=[current_time])],
# Phase 2.1 — bundle of plain Python callables exposed as one
# FastMCP namespace. The internal aggregator mounts it under
# ``/mcp/time`` on ``127.0.0.1:INTERNAL_MCP_PORT``; Phase 2.2's
# ClaudeCode adapter forwards that URL into
# ``BackendOptions.mcp_servers``. Phase 3's ``McpServerFrontend``
# reverse-proxies the same internal URL out to external clients.
McpServer.python_tool(name="time", tools=[current_time])
# Phase 3 — illustrates the ``lenient`` flag. Real-world stdio MCPs
# sometimes print "Processing..." or other chatter to stdout before
# their actual JSON-RPC frames; the default mcp client forwards
# those parse failures downstream as warnings (visible in
# Cursor/Cline). With ``lenient=True`` we silently drop non-JSON
# lines, so downstream UIs see clean JSON-RPC only. The command
# below is just a placeholder — replace with whatever stdio MCP
# you actually want gateway to ingest (e.g. an obsidian-mcp).
#
# Commented out by default: example users won't have the binary
# installed and an unreachable command makes ``docker compose up``
# surface a confusing "command not found" line at first request.
# Uncomment after pointing ``command`` at a real stdio MCP.
#
# McpServer.stdio(
# name="obsidian",
# command=["uvx", "mcp-obsidian"],
# env={"OBSIDIAN_API_KEY": "..."},
# lenient=True,
# ),
],
frontends=[ frontends=[
# Phase 1.4 — expose the agents as `model=<name>` on an
# Anthropic-compatible Messages endpoint. Auth comes from
# `BOOTSTRAP_TOKENS` in the env (`name1:value1,name2:value2`).
#
# Every HTTP frontend is mounted under its own path on the one
# gateway port (`Gateway.port`, 8000 here): `/anthropic/v1/messages`.
# Behind a reverse proxy set `Gateway(public_url="https://domain.com")`
# so advertised endpoints use the outside origin; the proxy just
# forwards everything to the gateway, no prefix stripping.
AnthropicMessagesFrontend(), AnthropicMessagesFrontend(),
# Phase 3 — re-exposes every declared `McpServer` outside the
# gateway with bearer auth + audit log. Each namespace lives at
# `/mcp/<name>/`; a flat bundle is published at `/mcp/all/`.
# Discovery page (HTML, auth-gated) at `/mcp/` with copy-pastable
# Cursor / Claude Desktop snippets. Auth re-uses `BOOTSTRAP_TOKENS`.
McpServerFrontend(), McpServerFrontend(),
# Phase 4.3 — browser admin UI. Creds come from
# `ADMIN_USER`/`ADMIN_PASS`; the session cookie is signed with
# `SESSION_SECRET`. Use it to mint tokens (Argon2-hashed in
# the DB), revoke them, and watch the audit log. Scope is
# enforced on the bearer frontends: tokens minted with scope
# `messages` only work on `/v1/messages`; `mcp` only on
# `/mcp/<name>`; `*` works everywhere. Served at `/admin/`; `/`
# redirects there.
AdminFrontend(), AdminFrontend(),
# Obsidian-vault chat frontend. Each `.md` is one conversation ApiFrontend(master_agent="assistant", branch_agent="assistant"),
# (User/Assistant turn pairs). The Obsidian companion plugin MarkdownFrontend(vault_path=VAULT / "chats", default_agent="assistant"),
# POSTs `{filename, content?}` to `/chat` — the frontend reads
# the file, runs the agent if the last turn is `user`, and
# appends the assistant reply back. With `log_all_chats=True`
# *every* turn (from Anthropic Messages too) is mirrored into
# `{vault}/_logs/<agent>/` so the vault is the central archive.
#
# `vault_path` here points at a per-restart tempdir so the
# example boots cleanly; in real deployments mount the
# Obsidian-sync container's vault volume to a stable path and
# pass that instead.
# Mounted at `/md` (`/md/chat`, `/md/chat/stream`).
MarkdownFrontend(
# Point at the dedicated chats subdir of your real Obsidian
# vault — the gateway has no idea (and no need) about other
# notes outside it. Path resolution / vault-escape checks
# are anchored here, so absolute-path attempts (and ``..``
# tricks) can't reach notes alongside it.
#
# vault_path=Path("/Users/me/Obsidian/Personal/chats"),
#
# Per-restart tempdir kept here so the example boots even
# without a real vault on the host.
vault_path=Path(tempfile.mkdtemp(prefix="beaver-vault-")).resolve(),
default_agent="research",
log_all_chats=True,
# ``chat_path`` (optional) overrides the default
# ``{vault}/_logs/<agent>/<date>_<slug>.md`` layout. Heads up:
# any custom path forces ``warm_index`` to scan the entire
# vault on startup so the fingerprint→file map of archived
# stateless chats survives a restart no matter where you put
# files.
chat_path=chat_path,
),
], ],
jobs=[Job("heartbeat", heartbeat, cron="0 * * * *")],
tz="UTC",
) )
+2 -3
View File
@@ -17,9 +17,8 @@ from pydantic import BaseModel, ConfigDict
class ExposedMcp: class ExposedMcp:
"""Reference to an ``McpServer`` (by name) exposed to a single agent. """Reference to an ``McpServer`` (by name) exposed to a single agent.
``tools`` is an allowlist of tool names (``None`` = every tool); ``tools`` allowlists tool names (``None`` = all); ``deny`` removes
``deny`` is a tuple of ``fnmatch`` patterns removed on top of that, ``fnmatch`` patterns on top, e.g. ``("delete_*",)``.
e.g. ``("delete_*",)``.
""" """
name: str name: str
+15 -21
View File
@@ -1,24 +1,20 @@
"""Claude agent definition, backed by the Claude Agent SDK. """Claude agent definition, backed by the Claude Agent SDK.
The system prompt is ``system_prompt`` verbatim or, per conversation kind, ``system_prompt`` is used verbatim, or per-kind granules from ``prompts``
the granules named in ``prompts`` assembled at every session spawn (see are assembled via ``agents/prompts.py``. ``skill_sets`` are
``core/prompt.py``). ``skill_sets`` are directories of ``<skill>/SKILL.md`` ``<skill>/SKILL.md`` directories, each becoming a local SDK plugin.
folders; each becomes a local SDK plugin, either the same tuple for every
kind or a ``SkillSets`` with a tuple per kind (§4.3: the master never sees
what a branch opens). Nothing from disk is loaded otherwise: the adapter
runs with ``setting_sources=[]``.
""" """
from __future__ import annotations from __future__ import annotations
from collections.abc import Mapping # noqa: TC003 - pydantic runtime from collections.abc import Mapping # noqa: TC003
from pathlib import Path # noqa: TC003 - pydantic runtime from pathlib import Path # noqa: TC003
from pydantic import BaseModel, ConfigDict, Field, model_validator from pydantic import BaseModel, ConfigDict, Field, model_validator
from beaver_gateway.agents.base import BaseAgent from beaver_gateway.agents.base import BaseAgent
from beaver_gateway.agents.policy import PolicyRule # noqa: TC001 - pydantic runtime from beaver_gateway.agents.policy import PolicyRule # noqa: TC001
from beaver_gateway.agents.prompts import PromptSource # noqa: TC001 - pydantic runtime from beaver_gateway.agents.prompts import PromptSource # noqa: TC001
from beaver_gateway.conversations.kinds import KINDS, Kind from beaver_gateway.conversations.kinds import KINDS, Kind
__all__ = ["ClaudeAgent", "ClaudeOptions", "Prompts", "SkillSets"] __all__ = ["ClaudeAgent", "ClaudeOptions", "Prompts", "SkillSets"]
@@ -50,10 +46,10 @@ class ClaudeOptions(BaseModel):
class Prompts(BaseModel): class Prompts(BaseModel):
"""Prompt assembly per conversation kind (§3.12): the granules, in order. """Prompt assembly per conversation kind: the granules, in order.
A kind left ``None`` is not served by the agent; ``ClaudeAgent.kinds`` A kind left ``None`` is not served; ``ClaudeAgent.kinds`` follows from
follows from the kinds set here. the kinds set here.
""" """
model_config = ConfigDict(frozen=True) model_config = ConfigDict(frozen=True)
@@ -94,14 +90,12 @@ class ClaudeAgent(BaseAgent):
prompts: Prompts = Field(default_factory=Prompts) prompts: Prompts = Field(default_factory=Prompts)
kinds: tuple[Kind, ...] = () kinds: tuple[Kind, ...] = ()
"""Conversation kinds this agent serves; ``create``/``spawn`` reject the """Conversation kinds this agent serves; ``create``/``spawn`` reject the rest.
rest. Defaults to the kinds ``prompts`` covers, or ``("deep",)`` for a Defaults to what ``prompts`` covers, or ``("deep",)`` for a verbatim prompt."""
verbatim ``system_prompt``."""
skill_sets: tuple[Path, ...] | SkillSets = () skill_sets: tuple[Path, ...] | SkillSets = ()
"""Skill-set directories, each a local plugin: one tuple for every kind, """Skill-set directories, each a local plugin: one tuple for every kind,
or ``SkillSets`` to give each kind its own (a kind left ``None`` gets or a ``SkillSets`` giving each kind its own (``None`` = no skills)."""
no skills)."""
gateway_tools: tuple[str, ...] = () gateway_tools: tuple[str, ...] = ()
"""Gateway tools exposed in-process (``read_conversation``, ``spawn``, """Gateway tools exposed in-process (``read_conversation``, ``spawn``,
@@ -109,8 +103,8 @@ class ClaudeAgent(BaseAgent):
options: ClaudeOptions = Field(default_factory=ClaudeOptions) options: ClaudeOptions = Field(default_factory=ClaudeOptions)
policy: tuple[PolicyRule, ...] = () policy: tuple[PolicyRule, ...] = ()
"""``PreToolUse`` rules (§3.7), run in order on every tool call; the """``PreToolUse`` rules, run in order on every tool call; the first
first ``Deny`` is what the model reads back. See ``core/policy``.""" ``Deny`` is what the model reads back. See ``agents/policy.py``."""
@model_validator(mode="after") @model_validator(mode="after")
def _kinds_follow_prompts(self) -> ClaudeAgent: def _kinds_follow_prompts(self) -> ClaudeAgent:
+10 -17
View File
@@ -1,18 +1,8 @@
"""PreToolUse policy - the boundary without permission prompts (§3.7). """PreToolUse policy - the boundary without permission prompts.
``bypassPermissions`` everywhere; what a model may do is decided by A rule is a callable ``(ToolCall) -> Deny | None`` declared per agent
mounts, ``disallowed_tools`` and the rules here. A rule is a callable (``ClaudeAgent.policy``); the first :class:`Deny` becomes the tool result
``(ToolCall) -> Deny | None`` declared per agent (``ClaudeAgent.policy``); the model reads back. A rule that raises is a deny too: fails closed.
the SDK backend registers one in-process ``PreToolUse`` hook that runs
the rules in order and turns the first :class:`Deny` into a hook deny
whose reason the model reads as the tool result. Rules never see
secrets and never prompt - they only say no, with a reason.
A rule that raises is a deny too: the boundary fails closed, the
traceback lands in the log.
Every tool call - allowed or denied - is reported to the audit sink the
backend was given, so the admin audit page shows what the model touched.
""" """
from __future__ import annotations from __future__ import annotations
@@ -90,11 +80,14 @@ class ToolCall:
return None return None
def resolve(self, raw: str) -> Path: def resolve(self, raw: str) -> Path:
"""Absolute path for ``raw``, normalised lexically.
Not via ``Path.resolve``, which would follow symlinks on the host,
not the model's view.
"""
path = Path(raw).expanduser() path = Path(raw).expanduser()
if not path.is_absolute(): if not path.is_absolute():
path = self.cwd / path path = self.cwd / path
# ``resolve`` would follow symlinks on the gateway host, which may
# not be the model's view; normalise lexically instead.
return Path(*_normalize(path.parts)) return Path(*_normalize(path.parts))
@@ -136,7 +129,7 @@ async def evaluate(
verdict = rule(call) verdict = rule(call)
if inspect.isawaitable(verdict): if inspect.isawaitable(verdict):
verdict = await verdict verdict = await verdict
except Exception: # noqa: BLE001 - a broken rule must fail closed except Exception: # noqa: BLE001
name = getattr(rule, "__name__", repr(rule)) name = getattr(rule, "__name__", repr(rule))
_log.exception("policy rule %s failed on %s", name, call.tool) _log.exception("policy rule %s failed on %s", name, call.tool)
return Deny(reason=f"policy rule {name} failed; the call is refused") return Deny(reason=f"policy rule {name} failed; the call is refused")
+3 -7
View File
@@ -1,12 +1,7 @@
"""System prompt assembly from a list of source files. """System prompt assembly from a list of source files.
The gateway holds no prompt text: an agent names its granules (paths from :func:`assemble` concatenates an agent's granules in order. A source is a
``config.py``) and :func:`assemble` concatenates them in that order, so the path, or a ``(tag, path)`` pair whose content is wrapped in ``<tag>...</tag>``.
result is byte-for-byte identical for every session of the same agent as
long as the files are. A source is a path, or a ``(tag, path)`` pair whose
content is wrapped in ``<tag>...</tag>`` - the markup lives here, the vault
keeps plain markdown. Each granule's hash is logged at assembly so a
drifted prompt can be traced to the file that changed.
""" """
from __future__ import annotations from __future__ import annotations
@@ -27,6 +22,7 @@ PromptSource = str | Path | tuple[str, str | Path]
def assemble(sources: Iterable[PromptSource]) -> str: def assemble(sources: Iterable[PromptSource]) -> str:
"""Concatenate ``sources`` in order into one system prompt."""
parts: list[str] = [] parts: list[str] = []
for source in sources: for source in sources:
tag, raw = source if isinstance(source, tuple) else (None, source) tag, raw = source if isinstance(source, tuple) else (None, source)
+12 -54
View File
@@ -1,31 +1,12 @@
"""Raycast agent definition. """Raycast agent definition.
Field set is the union of ``raycast_api.ChatAPI.stream`` parameters that Field set is the union of ``raycast_api.ChatAPI.stream`` parameters useful
make sense as **per-agent defaults**. Per-request values (currently: as per-agent defaults; per-request values override them where both exist.
``temperature``) win when both are set; the rest fall back to whatever
the agent declared, then to Raycast's own defaults.
``BaseAgent.system_prompt`` maps onto Raycast's wire field
``additional_system_instructions`` (the slot the real client uses for
*content*); the wire field ``system_instructions`` stays at the Raycast
source default — ``"markdown"`` for ``AI_CHAT``, ``"plain"`` otherwise.
We don't expose that wire dichotomy to the user — they get one
conceptual "system prompt".
Excluded on purpose:
* ``buffer_id``/``message_id``/``current_date`` — per-call ephemeral
* ``provider`` override — escape hatch for non-catalog models, no clear
use case yet (revisit in PRD §14 when discovery lands)
* ``locale`` — process-wide via ``Settings.raycast_locale`` because we
only spin up one ``raycast_api.Client`` per gateway
* ``system_instructions`` (wire) — that's a format marker, not content;
the SDK fills it from the source default and we let it
""" """
from __future__ import annotations from __future__ import annotations
from raycast_api import ( # noqa: F401 — UserPreferences re-exported for user configs from raycast_api import ( # noqa: F401
RemoteTool, RemoteTool,
Source, Source,
UserPreferences, UserPreferences,
@@ -36,44 +17,21 @@ from beaver_gateway.agents.base import BaseAgent
class RaycastAgent(BaseAgent): class RaycastAgent(BaseAgent):
"""Agent backed by ``raycast-api``. """Agent backed by ``raycast-api``."""
``available_native_tools`` is the closed set of Raycast's server-side
"remote tools" (``web_search``, ``search_images``, ``read_page``) —
typed as ``RemoteTool`` so config-time IDE completion lists exactly
the three valid values. Pydantic also coerces string literals, so
``("web_search", "read_page")`` keeps working unchanged.
``user_preferences`` toggles the auto-generated ``<user-preferences>``
block Raycast prepends to ``additional_system_instructions``:
* ``True`` (default) → auto from host locale/timezone/today, rebuilt
every request so the date stays fresh;
* ``False`` → omit the block entirely;
* ``UserPreferences(...)`` instance → used verbatim (frozen at the
time the agent was loaded, so the date won't auto-update);
* ``Callable[[], UserPreferencesArg]`` → re-invoked on every request.
Use this for the common case "fresh date but custom
locale/timezone": ``user_preferences=lambda:
UserPreferences(locale="ru-RU", timezone="Europe/Berlin",
current_date=date.today().isoformat())``. Callables may nest
(a lambda returning a lambda…) but there's no real reason to.
The library uses this block for date/locale-aware formatting, not
for personalisation/memory — those are out of scope upstream.
``reasoning_effort`` values vary by model: GPT-5 takes
``"minimal"|"low"|"medium"|"high"``; Anthropic exposes nothing here
(Claude reasoning lives in a separate ``…-reasoning`` model variant
in the catalog). Unknown effort for the chosen model is ignored
server-side, so we stay loose as ``str | None``.
"""
streaming: bool = True streaming: bool = True
available_native_tools: tuple[RemoteTool, ...] = () available_native_tools: tuple[RemoteTool, ...] = ()
"""Raycast's server-side "remote tools": ``web_search``, ``search_images``,
``read_page``."""
source: Source = Source.AI_CHAT source: Source = Source.AI_CHAT
temperature: float | None = None temperature: float | None = None
reasoning_effort: str | None = None reasoning_effort: str | None = None
"""Model-specific (e.g. GPT-5: ``"minimal"``/``"low"``/``"medium"``/``"high"``);
ignored server-side if the model doesn't support it."""
tool_choice: str | None = None tool_choice: str | None = None
user_preferences: UserPreferencesArg = True user_preferences: UserPreferencesArg = True
"""Auto-fills the ``<user-preferences>`` block from host locale/timezone/date;
``False`` omits it, or pass ``UserPreferences(...)`` / a zero-arg callable."""
+10 -13
View File
@@ -2,18 +2,7 @@
A backend turns an Anthropic-style turn (``messages`` + agent definition) A backend turns an Anthropic-style turn (``messages`` + agent definition)
into a stream of :class:`~beaver_gateway.events.stream.MessageStreamEvent` into a stream of :class:`~beaver_gateway.events.stream.MessageStreamEvent`
records. The frontend serializes whatever comes out straight to SSE, so records; provider quirks are translated here, not in the frontend.
backends are the only place where provider quirks are translated.
Implementations are plain :class:`typing.Protocol` conformers - no ABC
subclassing - to keep them swappable in tests with bare async generators.
``**options`` is the one extension point. Known keys, all optional and
ignored by backends that don't keep state: ``conversation_id`` (stable id
the backend may pin a live session to), ``session_id`` (backend session to
resume when nothing is live), ``capture`` (a
:class:`~beaver_gateway.backends.capture.TurnCapture` the backend fills
after the stream closes).
""" """
from __future__ import annotations from __future__ import annotations
@@ -40,5 +29,13 @@ class Backend(Protocol):
system: str | None = None, system: str | None = None,
**options: Any, **options: Any,
) -> AsyncIterator[MessageStreamEvent]: ) -> AsyncIterator[MessageStreamEvent]:
"""Yield Anthropic stream events for one turn against ``agent``.""" """Yield Anthropic stream events for one turn against ``agent``.
``**options`` keys are optional and ignored by backends that don't
keep state: ``conversation_id`` (stable id the backend may pin a
live session to), ``session_id`` (backend session to resume when
nothing is live), ``capture`` (a
:class:`~beaver_gateway.backends.capture.TurnCapture` the backend
fills after the stream closes).
"""
... ...
+1 -4
View File
@@ -1,10 +1,7 @@
"""Side channel for what a backend learned during one turn. """Side channel for what a backend learned during one turn.
Frontends pass ``capture=TurnCapture()`` through ``Backend.complete``'s Frontends pass ``capture=TurnCapture()`` through ``Backend.complete``'s
``**options``. Backends that keep state per conversation (the Claude SDK ``**options``; only backends that keep state per conversation fill it in.
adapter) fill it in after the stream closes; backends that don't
(anthropic, raycast) drop the kwarg and the frontend falls back to a
text-only history.
""" """
from __future__ import annotations from __future__ import annotations
+31 -35
View File
@@ -1,29 +1,7 @@
"""Claude Agent SDK backend adapter. """Claude Agent SDK backend adapter.
One :class:`ClaudeSdkBackend` per :class:`ClaudeAgent`. A live session is One :class:`ClaudeSdkBackend` per :class:`ClaudeAgent`; its sessions live in
one ``ClaudeSDKClient`` (one claude subprocess) and runs one turn at a a shared :class:`~beaver_gateway.backends.sessions.SessionPool`.
time; the sessions of every agent live in one shared
:class:`~beaver_gateway.backends.sessions.SessionPool` that owns TTL and
memory-pressure eviction. Sessions are keyed by ``conversation_id`` when
the caller passes one or by a text-only fingerprint of ``messages[:-1]``
for stateless callers (``/v1/messages``). Without a live session the
adapter resumes ``session_id`` from the session store (after closing any
``tool_use`` left open by a crash), or seeds the incoming history into the
store via ``core/transcript`` and resumes that.
Per-turn ``**options`` beyond the protocol's: ``kind`` (conversation kind,
picks the prompt assembly and the pool TTL), ``pinned`` (never evicted),
``tools=False`` (no MCP at all - forks and jobs), ``observer`` (callback
receiving every raw SDK message, subagent ones included), ``turn_id``.
Events on the wire are the Anthropic ``MessageStreamEvent`` family: one
``message_start``/``message_stop`` envelope per turn, block indices
rebased across the API calls claude makes inside the turn.
Process isolation: claude is spawned through a small exec wrapper that
drops every inherited environment variable outside a whitelist and, when
``RunnerConfig.user`` is set, switches to that uid before exec (done in the
wrapper rather than via ``subprocess(user=...)``, which uvloop rejects).
""" """
from __future__ import annotations from __future__ import annotations
@@ -113,8 +91,6 @@ if TYPE_CHECKING:
_log = logging.getLogger("beaver_gateway.backends.claude_sdk") _log = logging.getLogger("beaver_gateway.backends.claude_sdk")
# §3.7: in bypass the callback only ever sees AskUserQuestion, and that is
# exactly the one we want - the SDK's warning about the rest is noise here.
warnings.filterwarnings("ignore", category=CanUseToolShadowedWarning) warnings.filterwarnings("ignore", category=CanUseToolShadowedWarning)
ASK_TOOL = "AskUserQuestion" ASK_TOOL = "AskUserQuestion"
@@ -176,8 +152,7 @@ AuditSink = "Callable[[policy_mod.ToolAudit], Awaitable[None]]"
Asker = "Callable[[str, dict[str, Any]], Awaitable[str]]" Asker = "Callable[[str, dict[str, Any]], Awaitable[str]]"
"""``(conversation_key, AskUserQuestion input) -> text the model reads as the """``(conversation_key, AskUserQuestion input) -> text the model reads as the
tool result``. The only channel an answer has in bypass mode is tool result``. The only channel an answer has in bypass mode is
``PermissionResultDeny.message`` (spike S1, s05): ``updated_input`` never ``PermissionResultDeny.message``: ``updated_input`` never reaches the model."""
reaches the model."""
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
@@ -306,7 +281,7 @@ class ClaudeSdkBackend:
*, *,
agent: BaseAgent, agent: BaseAgent,
messages: Iterable[MessageParam], messages: Iterable[MessageParam],
system: str | None = None, # noqa: ARG002 - the agent owns its prompt system: str | None = None, # noqa: ARG002
conversation_id: str | None = None, conversation_id: str | None = None,
session_id: str | None = None, session_id: str | None = None,
reseed: bool = False, reseed: bool = False,
@@ -316,8 +291,21 @@ class ClaudeSdkBackend:
tools: bool = True, tools: bool = True,
observer: Callable[[Any], None] | None = None, observer: Callable[[Any], None] | None = None,
turn_id: str | None = None, turn_id: str | None = None,
**options: Any, # noqa: ARG002 - per-request knobs are not supported **options: Any, # noqa: ARG002
) -> AsyncIterator[MessageStreamEvent]: ) -> AsyncIterator[MessageStreamEvent]:
"""Run one turn against ``agent``, yielding Anthropic stream events.
Sessions are keyed by ``conversation_id`` when given, else a
text-only fingerprint of ``messages[:-1]``. Without a live session,
resumes ``session_id`` (closing any ``tool_use`` a crash left open),
or seeds ``messages[:-1]`` via ``backends/transcript.py`` and resumes
that. ``kind`` picks the prompt assembly and pool TTL; ``pinned``
sessions are never evicted; ``tools=False`` mounts no MCP servers
(forks and jobs); ``observer`` sees every raw SDK message.
A resumed session that dies before any event reached the caller is
reseeded from history and retried once.
"""
if agent.name != self._agent.name: if agent.name != self._agent.name:
msg = f"backend bound to {self._agent.name!r}, got {agent.name!r}" msg = f"backend bound to {self._agent.name!r}, got {agent.name!r}"
raise ValueError(msg) raise ValueError(msg)
@@ -339,8 +327,6 @@ class ClaudeSdkBackend:
live.running_turn = turn_id or message_id live.running_turn = turn_id or message_id
live.last_used = time.monotonic() live.last_used = time.monotonic()
try: try:
# Events go out as the CLI produces them: the frontends
# stream text and thinking live, the turn is not buffered.
async for event in self._run_turn( async for event in self._run_turn(
live, prompt, turn, observer, capture live, prompt, turn, observer, capture
): ):
@@ -348,8 +334,6 @@ class ClaudeSdkBackend:
except Exception: except Exception:
live.running_turn = None live.running_turn = None
await self._pool.close(key) await self._pool.close(key)
# A dead resume can be reseeded from history, but only
# while nothing of this turn has reached the caller yet.
if not (live.resumed and live.turns == 0) or turn.events: if not (live.resumed and live.turns == 0) or turn.events:
raise raise
_log.exception( _log.exception(
@@ -626,6 +610,12 @@ class ClaudeSdkBackend:
) -> Callable[ ) -> Callable[
[str, dict[str, Any], ToolPermissionContext], Awaitable[PermissionResult] [str, dict[str, Any], ToolPermissionContext], Awaitable[PermissionResult]
]: ]:
"""Build the ``can_use_tool`` callback for one session.
In bypass mode the SDK only ever calls this back for
``AskUserQuestion``; the warning about the rest is suppressed at
import time.
"""
asker = self._asker asker = self._asker
async def can_use_tool( async def can_use_tool(
@@ -648,7 +638,7 @@ class ClaudeSdkBackend:
def _hooks( def _hooks(
self, key: str, spec: _SessionSpec self, key: str, spec: _SessionSpec
) -> dict[HookEvent, list[HookMatcher]] | None: ) -> dict[HookEvent, list[HookMatcher]] | None:
"""§3.7: one in-process ``PreToolUse`` hook - policy rules, then audit.""" """One in-process ``PreToolUse`` hook: policy rules, then audit."""
agent = self._agent agent = self._agent
if not agent.policy and self._audit_sink is None: if not agent.policy and self._audit_sink is None:
return None return None
@@ -708,6 +698,12 @@ class ClaudeSdkBackend:
return plugins return plugins
def _exec_wrapper(self, *, extra_keep: tuple[str, ...]) -> Path: def _exec_wrapper(self, *, extra_keep: tuple[str, ...]) -> Path:
"""Path to a generated wrapper that execs claude with a whitelisted env.
Drops every inherited variable outside the whitelist and, when
``RunnerConfig.user`` is set, switches to that uid before exec - done
here rather than via ``subprocess(user=...)``, which uvloop rejects.
"""
if self._wrapper is not None: if self._wrapper is not None:
return self._wrapper return self._wrapper
keep = sorted({*ENV_KEEP, *self._agent.options.env_keep, *extra_keep}) keep = sorted({*ENV_KEEP, *self._agent.options.env_keep, *extra_keep})
+31 -87
View File
@@ -1,32 +1,8 @@
"""Raycast backend adapter. """Raycast backend adapter.
Translates between Anthropic's ``/v1/messages`` wire vocabulary (incoming Translates between Anthropic's ``/v1/messages`` wire vocabulary and the
``MessageParam`` history, outgoing ``MessageStreamEvent`` SSE) and the ``raycast-api`` SDK, splicing MCP tools into Raycast's ``tool_calls`` and
``raycast-api`` SDK (``Message`` history, ``ChatStreamChunk`` SSE). routing them back in-process where the agent declares ``expose_mcps``.
Two halves live here:
* :func:`_to_raycast_messages` — pure conversion of an Anthropic message
list into ``list[raycast_api.Message]``. ``tool_result`` blocks carry no
tool name in Anthropic; we recover it by remembering each ``tool_use``
id we saw upstream.
* :meth:`RaycastBackend.complete` — opens a ``client.chat.stream`` and
walks chunks through a tiny block-state machine. The state machine
exists only because Raycast streams ``tool_calls`` in three phases
(open with id+name, deltas with empty id, final summary with the full
``arguments``) — Anthropic wants one ``content_block_start`` → deltas
→ ``content_block_stop`` per block, so we de-duplicate the final
summary against the per-delta increments already emitted.
MCP wiring: Raycast has no native MCP concept, so when an agent declares
``expose_mcps`` we splice each MCP's tools into the wire request as
``Tool.local(name=f"{mcp}__{tool}", ...)`` and run a gateway-internal
loop. Every time the model emits a tool_call for one of those local
tools we route it back to the underlying MCP in-process, append the
result as a ``tool`` message, and re-issue the stream — all inside one
Anthropic envelope (one ``message_start`` … one ``message_stop``). The
tool_use blocks DO surface to the caller (mirrors what
``ClaudeSdkBackend`` does), but tool_results stay internal.
""" """
from __future__ import annotations from __future__ import annotations
@@ -76,11 +52,9 @@ __all__ = ["RaycastBackend"]
_log = logging.getLogger("beaver_gateway.backends.raycast") _log = logging.getLogger("beaver_gateway.backends.raycast")
# Cap on consecutive tool-call turns inside one Anthropic envelope.
# Real conversations rarely chain more than a handful; the limit only
# fires on a model that loops, and surfaces as a clean ``end_turn``
# with an error tool_result rather than a hang.
_MAX_TOOL_TURNS = 20 _MAX_TOOL_TURNS = 20
"""Cap on consecutive tool-call turns inside one Anthropic envelope; a model
that loops hits this and gets a clean ``end_turn`` instead of a hang."""
_RAYCAST_TO_ANTHROPIC_STOP: dict[str, StopReason] = { _RAYCAST_TO_ANTHROPIC_STOP: dict[str, StopReason] = {
@@ -392,7 +366,7 @@ class _BlockState:
def __init__(self) -> None: def __init__(self) -> None:
self.index: int = -1 self.index: int = -1
self.kind: str | None = None # "text" | "thinking" | "tool_use" | None self.kind: str | None = None
self.tool_id_to_block: dict[str, int] = {} self.tool_id_to_block: dict[str, int] = {}
self.tool_idx_to_id: dict[int, str] = {} self.tool_idx_to_id: dict[int, str] = {}
@@ -428,11 +402,14 @@ class RaycastBackend:
self._client = client self._client = client
self._mcp_servers: Mapping[str, FastMCP] = mcp_servers or {} self._mcp_servers: Mapping[str, FastMCP] = mcp_servers or {}
self._mcp_tools: Mapping[str, list[FastMCPTool]] = mcp_tools or {} self._mcp_tools: Mapping[str, list[FastMCPTool]] = mcp_tools or {}
# Cached per-agent catalog — agents are immutable, so a single
# render at first use covers the gateway's lifetime.
self._agent_catalog: dict[str, _AgentToolCatalog] = {} self._agent_catalog: dict[str, _AgentToolCatalog] = {}
def _catalog_for(self, agent: RaycastAgent) -> _AgentToolCatalog: def _catalog_for(self, agent: RaycastAgent) -> _AgentToolCatalog:
"""Build (or reuse) the tool catalog for ``agent``.
Agents are immutable, so a single render at first use covers the
gateway's lifetime.
"""
cached = self._agent_catalog.get(agent.name) cached = self._agent_catalog.get(agent.name)
if cached is not None: if cached is not None:
return cached return cached
@@ -448,31 +425,26 @@ class RaycastBackend:
system: str | None = None, system: str | None = None,
**options: Any, **options: Any,
) -> AsyncIterator[MessageStreamEvent]: ) -> AsyncIterator[MessageStreamEvent]:
"""Run one turn against ``agent``, yielding Anthropic stream events.
On the wire Raycast uses ``system_instructions`` as a format marker
(``"markdown"``/``"plain"``, filled from the source default when we
pass ``None``) and ``additional_system_instructions`` for the actual
prompt content; ``system`` (or ``agent.system_prompt``) flows into
the latter.
"""
if not isinstance(agent, RaycastAgent): if not isinstance(agent, RaycastAgent):
msg = f"RaycastBackend requires RaycastAgent, got {type(agent).__name__}" msg = f"RaycastBackend requires RaycastAgent, got {type(agent).__name__}"
raise TypeError(msg) raise TypeError(msg)
raycast_messages = _to_raycast_messages(messages) raycast_messages = _to_raycast_messages(messages)
catalog = self._catalog_for(agent) catalog = self._catalog_for(agent)
# Native remote tools + spliced MCP locals. ``None`` keeps the
# SDK from sending a ``tools`` field at all when neither is
# declared.
tools_arg: list[Tool | RemoteTool | str] | None = ( tools_arg: list[Tool | RemoteTool | str] | None = (
list(catalog.tools) if catalog.tools else None list(catalog.tools) if catalog.tools else None
) )
# On the wire Raycast uses ``system_instructions`` as a format
# marker (``"markdown"`` for AI_CHAT, ``"plain"`` otherwise —
# filled in by the SDK from the source default when we pass
# ``None``) and ``additional_system_instructions`` as the actual
# prompt content. So our ``system_prompt`` (or the per-request
# Anthropic ``system``, if present) flows into the *additional*
# slot. The SDK still prepends ``<user-preferences>`` to whatever
# we hand it via ``_build_preamble``.
prompt_content = system if system is not None else agent.system_prompt prompt_content = system if system is not None else agent.system_prompt
# Per-request options win over agent defaults; agent defaults
# win over Raycast SDK defaults. ``None`` means "fall back".
async for event in self._stream( async for event in self._stream(
agent=agent, agent=agent,
raycast_messages=raycast_messages, raycast_messages=raycast_messages,
@@ -499,12 +471,16 @@ class RaycastBackend:
reasoning_effort: str | None, reasoning_effort: str | None,
tool_choice: str | None, tool_choice: str | None,
) -> AsyncIterator[MessageStreamEvent]: ) -> AsyncIterator[MessageStreamEvent]:
"""Yield one Anthropic envelope spanning one or more ``chat.stream`` calls.
A tool_call not in ``mcp_routing`` is left to bubble out as a
regular ``tool_use`` block instead, for the caller to answer with a
``tool_result`` the usual way.
"""
message_id = f"msg_{uuid.uuid4().hex}" message_id = f"msg_{uuid.uuid4().hex}"
yield build_message_start(message_id=message_id, model=agent.model) yield build_message_start(message_id=message_id, model=agent.model)
state = _BlockState() state = _BlockState()
# ``working_messages`` is the rolling history fed back to
# Raycast as we resolve MCP tool_calls turn by turn.
working_messages = list(raycast_messages) working_messages = list(raycast_messages)
last_usage: dict[str, int] | None = None last_usage: dict[str, int] | None = None
last_finish: str | None = None last_finish: str | None = None
@@ -517,9 +493,6 @@ class RaycastBackend:
model=agent.model, model=agent.model,
messages=working_messages, messages=working_messages,
source=agent.source, source=agent.source,
# ``system_instructions=None`` → SDK substitutes the
# source default (``"markdown"`` / ``"plain"``). Real
# prompt goes into ``additional_system_instructions``.
additional_system_instructions=prompt_content, additional_system_instructions=prompt_content,
user_preferences=agent.user_preferences, user_preferences=agent.user_preferences,
tools=tools, tools=tools,
@@ -540,27 +513,16 @@ class RaycastBackend:
if acc.usage: if acc.usage:
last_usage = acc.usage last_usage = acc.usage
# Figure out which (if any) tool_calls land on our MCP
# routing table. Anything not in the table is left to bubble
# out of the envelope as a regular tool_use block — the
# Anthropic caller can then respond with a tool_result the
# usual way, and the next ``complete`` invocation will
# carry it back in.
pending_mcp_calls = [ pending_mcp_calls = [
tid for tid in acc.tool_order if acc.tool_names.get(tid) in mcp_routing tid for tid in acc.tool_order if acc.tool_names.get(tid) in mcp_routing
] ]
if not pending_mcp_calls: if not pending_mcp_calls:
break break
# Close whatever block is still open before we step into
# tool execution — the next turn's chunks start a fresh
# block sequence.
if state.kind is not None: if state.kind is not None:
yield build_content_block_stop(state.index) yield build_content_block_stop(state.index)
state.kind = None state.kind = None
# Echo the assistant turn so Raycast sees its own reply +
# tool_calls in subsequent context.
assistant_text = "".join(acc.text_parts) assistant_text = "".join(acc.text_parts)
assistant_tool_calls = [ assistant_tool_calls = [
ToolCall( ToolCall(
@@ -576,10 +538,6 @@ class RaycastBackend:
) )
) )
# Dispatch each MCP-routed call and append a ``tool``
# message. Calls not in the routing table get a placeholder
# error so the model can correct itself rather than the
# gateway hanging the conversation.
for tid in acc.tool_order: for tid in acc.tool_order:
tool_name = acc.tool_names.get(tid, "") tool_name = acc.tool_names.get(tid, "")
args_str = "".join(acc.tool_args.get(tid, [])) or "{}" args_str = "".join(acc.tool_args.get(tid, [])) or "{}"
@@ -588,16 +546,12 @@ class RaycastBackend:
tool_name=tool_name, args_json=args_str, mcp_routing=mcp_routing tool_name=tool_name, args_json=args_str, mcp_routing=mcp_routing
) )
else: else:
# Non-MCP tool — shouldn't really happen because we
# haven't surfaced any other locals, but defend.
result_text = f"Tool {tool_name!r} is not handled by the gateway." result_text = f"Tool {tool_name!r} is not handled by the gateway."
working_messages.append( working_messages.append(
RaycastMessage.tool( RaycastMessage.tool(
tool_call_id=tid, name=tool_name, result=result_text tool_call_id=tid, name=tool_name, result=result_text
) )
) )
# Loop: re-stream with the new history.
else: else:
_log.warning( _log.warning(
"raycast tool-call loop hit %d-turn cap for agent %r; " "raycast tool-call loop hit %d-turn cap for agent %r; "
@@ -606,7 +560,6 @@ class RaycastBackend:
agent.name, agent.name,
) )
# Close whatever block is still open before the final delta.
if state.kind is not None: if state.kind is not None:
yield build_content_block_stop(state.index) yield build_content_block_stop(state.index)
state.kind = None state.kind = None
@@ -716,6 +669,12 @@ class RaycastBackend:
the final-summary chunk's arguments string is dropped because the the final-summary chunk's arguments string is dropped because the
deltas already streamed it. deltas already streamed it.
Providers disagree on how ``arguments`` arrives: streaming ones
(GPT) send deltas across chunks and then restate the full string in
a final summary chunk (skip it, already streamed); non-streaming
ones (Gemini) send it only in that final summary (emit it once,
since nothing streamed it first).
Side-effects on ``acc`` mirror what gets emitted to the wire so Side-effects on ``acc`` mirror what gets emitted to the wire so
the gateway can rebuild a full ``ToolCall`` for the next Raycast the gateway can rebuild a full ``ToolCall`` for the next Raycast
turn (it needs the joined ``arguments`` JSON string, which the turn (it needs the joined ``arguments`` JSON string, which the
@@ -728,9 +687,6 @@ class RaycastBackend:
raw_tc = raw_tcs[i] if i < len(raw_tcs) else {} raw_tc = raw_tcs[i] if i < len(raw_tcs) else {}
idx_field = raw_tc.get("index") if isinstance(raw_tc, dict) else None idx_field = raw_tc.get("index") if isinstance(raw_tc, dict) else None
# Resolve this entry to a tool-id key, mirroring
# `raycast_api.ChatResult._merge_tool_calls`. Phase 1 carries
# id+index, phase 2 only index, phase 3 only id.
tool_id: str | None = None tool_id: str | None = None
if tc.id: if tc.id:
tool_id = tc.id tool_id = tc.id
@@ -743,7 +699,6 @@ class RaycastBackend:
block_idx = state.tool_id_to_block.get(tool_id) block_idx = state.tool_id_to_block.get(tool_id)
if block_idx is None: if block_idx is None:
# New tool_use block. Close any open text/thinking block first.
if state.kind is not None: if state.kind is not None:
events.append(build_content_block_stop(state.index)) events.append(build_content_block_stop(state.index))
state.index += 1 state.index += 1
@@ -756,22 +711,11 @@ class RaycastBackend:
block_idx, tool_use_id=tool_id, name=tc.name or "" block_idx, tool_use_id=tool_id, name=tc.name or ""
) )
) )
# Streaming providers (GPT) deliver arguments as deltas
# across chunks and re-state the full string in a final
# summary chunk; non-streaming-args providers (Gemini)
# only send args in the final summary. Emit args in
# both cases when this is the first appearance — final
# summary then IS the full string.
if tc.arguments: if tc.arguments:
events.append(build_input_json_delta(block_idx, tc.arguments)) events.append(build_input_json_delta(block_idx, tc.arguments))
acc.add_tool_args(tool_id, tc.arguments) acc.add_tool_args(tool_id, tc.arguments)
continue continue
# Existing block. Final summary chunks restate the full args
# string; if we already streamed deltas, that restatement is
# a duplicate (skip). If we streamed nothing (the streaming
# provider didn't send mid-arg deltas — Gemini path again),
# the summary IS the args — emit it once.
if is_final_summary: if is_final_summary:
if tc.arguments and not acc.tool_args.get(tool_id): if tc.arguments and not acc.tool_args.get(tool_id):
events.append(build_input_json_delta(block_idx, tc.arguments)) events.append(build_input_json_delta(block_idx, tc.arguments))
+11 -10
View File
@@ -1,13 +1,9 @@
"""Pool of live Agent SDK sessions across every Claude agent (§3.2). """Pool of live Agent SDK sessions across every Claude agent.
One :class:`Session` is one ``ClaudeSDKClient`` (one claude subprocess). One :class:`Session` is one ``ClaudeSDKClient`` (one claude subprocess). The
The pool owns the two decisions the adapters used to make on their own: pool decides when a session is closed for idleness (TTL by conversation
when a session is closed for idleness (TTL by conversation kind) and kind) and which one goes when memory runs out (RSS of the subprocess tree
which one goes when memory runs out (measured RSS of the subprocess tree against the cgroup limit, or ``max_live`` with no limit).
against the cgroup limit, ``max_live`` where there is no limit). Eviction
only ever picks ``idle && !running_turn && !pending_question`` sessions
that are neither pinned (the master) nor ``dirty`` (mirror gap not yet
repaired); forks and jobs go first.
""" """
from __future__ import annotations from __future__ import annotations
@@ -69,7 +65,7 @@ class Session:
created_at: float = field(default_factory=time.monotonic) created_at: float = field(default_factory=time.monotonic)
turns: int = 0 turns: int = 0
state: dict[str, Any] = field(default_factory=dict) state: dict[str, Any] = field(default_factory=dict)
"""Scratch for policy rules (``core/policy``); dies with the process.""" """Scratch for policy rules (``agents/policy.py``); dies with the process."""
@property @property
def busy(self) -> bool: def busy(self) -> bool:
@@ -77,6 +73,10 @@ class Session:
@property @property
def evictable(self) -> bool: def evictable(self) -> bool:
"""Never evict a pinned (master) session or a ``dirty`` one.
``dirty`` means a mirror gap not yet repaired.
"""
return not (self.pinned or self.dirty or self.busy or self.pending_question) return not (self.pinned or self.dirty or self.busy or self.pending_question)
@property @property
@@ -167,6 +167,7 @@ class SessionPool:
return len(self._sessions) >= self._max_live return len(self._sessions) >= self._max_live
def victims(self) -> list[Session]: def victims(self) -> list[Session]:
"""Evictable sessions, forks/jobs first, then oldest idle."""
candidates = [s for s in self._sessions.values() if s.evictable] candidates = [s for s in self._sessions.values() if s.evictable]
candidates.sort(key=lambda s: (_EVICT_ORDER.get(s.kind, 1), s.last_used)) candidates.sort(key=lambda s: (_EVICT_ORDER.get(s.kind, 1), s.last_used))
return candidates return candidates
+10 -13
View File
@@ -1,16 +1,7 @@
"""Anthropic messages <-> Agent SDK transcript entries. """Anthropic messages <-> Agent SDK transcript entries.
:func:`build_entries` renders a message list as the entries the Claude CLI :func:`build_entries` renders a message list as entries the Claude CLI
itself writes (reference: ``t/spike_sdk/entries_reference.json``, CLI itself writes; :func:`messages_from_entries` is the projection back.
2.1.248 via ``import_session_to_store``): one ``user`` entry per prompt,
one ``assistant`` entry per content block sharing a message id, one
``user`` entry per ``tool_result`` parented on the matching ``tool_use``
entry. Only ``user``/``assistant`` entries are produced - no attachments,
titles or queue markers. Appending the result to a session store and
resuming that session id seeds an external history into the SDK.
:func:`messages_from_entries` is the projection back, used by tests and by
anything that needs Anthropic-shape history out of a mirrored transcript.
""" """
from __future__ import annotations from __future__ import annotations
@@ -61,6 +52,14 @@ def build_entries(
permission_mode: str = "bypassPermissions", permission_mode: str = "bypassPermissions",
now: datetime | None = None, now: datetime | None = None,
) -> list[dict[str, Any]]: ) -> list[dict[str, Any]]:
"""Render ``messages`` as entries a session store can resume from.
One ``user`` entry per prompt, one ``assistant`` entry per content
block sharing a message id, one ``user`` entry per ``tool_result``
parented on the matching ``tool_use``. Appending the result to a
session store and resuming that session id seeds external history
into the SDK.
"""
stamp = (now or datetime.now(UTC)).strftime("%Y-%m-%dT%H:%M:%S.") + ( stamp = (now or datetime.now(UTC)).strftime("%Y-%m-%dT%H:%M:%S.") + (
f"{(now or datetime.now(UTC)).microsecond // 1000:03d}Z" f"{(now or datetime.now(UTC)).microsecond // 1000:03d}Z"
) )
@@ -293,8 +292,6 @@ def _zero_usage() -> dict[str, Any]:
} }
# ---- repair, windows, projections ---------------------------------------
_PROMPT_TYPES = ("user", "assistant") _PROMPT_TYPES = ("user", "assistant")
_INTERRUPTED = "interrupted" _INTERRUPTED = "interrupted"
+5 -10
View File
@@ -1,13 +1,8 @@
"""Closing a deep chat: the digest, the index, the file line cap (§6.4, §8.4). """Closing a deep chat: the digest, the index, the file line cap.
The gateway knows no path by itself (§0.8): ``Distiller`` from ``config.py`` ``Distiller`` says where digests and the index live; the gateway checks the
names the agent, says where digests land and where the index lives, and digest's frontmatter and caps the merge text. ``LineCap`` is the same idea
the distiller writes the file on its own. What the gateway does is check that a file for a file a job rewrites: too long, and it is bounced back unchanged.
with a valid frontmatter appeared under ``Distiller.dir`` during the fork
turn, put one line into the index, and cap the merge text at
``SUMMARY_LINES``. ``LineCap`` is the same idea for a file a job rewrites
(``состояние.md``): a result longer than the cap is bounced - the file
goes back to what it was and the job is told to shorten.
""" """
from __future__ import annotations from __future__ import annotations
@@ -39,7 +34,7 @@ __all__ = [
] ]
SUMMARY_LINES = 5 SUMMARY_LINES = 5
"""A merge into the master is at most this many lines (§6.4).""" """A merge into the master is at most this many lines."""
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
+4 -9
View File
@@ -1,13 +1,8 @@
"""Persisted per-conversation queue, ``urgent > user > wake > normal`` (§3.4). """Persisted per-conversation queue: ``urgent > user > wake > normal``.
One ``ClaudeSDKClient`` runs one turn at a time, so ordering has to happen conversations/service.py runs one worker per conversation over these rows.
before the client: the rows here are the queue, ``core/conversations`` A row still ``running`` when the gateway starts was cut by a restart and is
runs one worker per conversation over them. ``urgent`` cuts a running flagged ``interrupted``, never re-run.
turn, ``user`` is the human, ``wake`` starts a turn as soon as the
conversation is idle and takes the queued normals with it, ``normal``
waits for the batching window or rides with the next turn. A row that is
still ``running`` when the gateway starts was cut by a restart; it is
flagged ``interrupted`` and never re-run.
""" """
from __future__ import annotations from __future__ import annotations
+1 -1
View File
@@ -1,4 +1,4 @@
"""Conversation kinds (§3.1) as one closed type for agents, frontends, service.""" """Conversation kinds as one closed type for agents, frontends, service."""
from __future__ import annotations from __future__ import annotations
+3 -7
View File
@@ -1,11 +1,7 @@
"""Master rotation (§4.5, §8.1, §8.3). """Master rotation: one logical master thread, many physical sessions.
One logical master thread, many physical sessions: when the policy says When the policy says so, a new master takes over the window; the old one
so, a new master is spawned and takes over the window atomically, the old hands out, closes, and its branches and queued normal injects move over.
one writes its handout as its last turn, closes, its finished branches get
marked in their windows, its queued normal injects move over, and the new
one receives "new day". Silence is measured by the user's messages only -
injects never extend a day.
""" """
from __future__ import annotations from __future__ import annotations
+6 -8
View File
@@ -1,8 +1,7 @@
"""In-process MCP server with the gateway's own tools (§3.1, §3.2). """In-process MCP server with the gateway's own tools.
One server per live session so every tool knows which conversation is One server per live session so every tool knows which conversation is
calling; ``alwaysLoad`` keeps the tools out of tool search. Which names a calling. Which names a session gets comes from ``ClaudeAgent.gateway_tools``.
session gets comes from ``ClaudeAgent.gateway_tools``.
""" """
from __future__ import annotations from __future__ import annotations
@@ -68,11 +67,10 @@ def build_tool_server(
def _redacting(spec: SdkMcpTool[Any]) -> SdkMcpTool[Any]: def _redacting(spec: SdkMcpTool[Any]) -> SdkMcpTool[Any]:
"""Put a tool's result through the same mask as every other MCP. """Put a tool's result through the same mask as every other MCP.
These tools are mounted in-process by the SDK, so they bypass the These tools are mounted in-process, bypassing the ``FastMCP`` middleware
``FastMCP`` middleware in :mod:`beaver_gateway.mcp.redacting` and in :mod:`beaver_gateway.mcp.redacting`, so the filter is attached here
need the filter attached here instead. ``read_conversation`` is the instead. ``read_conversation`` earns it: a replayed transcript can still
one that earns it: it replays a transcript, and a transcript written hold a credential written before redaction existed.
before any of this existed can still hold a credential.
""" """
inner = spec.handler inner = spec.handler
+1
View File
@@ -0,0 +1 @@
"""The in-process event bus and the stream event protocol backends emit."""
+2 -5
View File
@@ -1,10 +1,7 @@
"""In-process event bus: what the gateway does, as a stream frontends can tap. """In-process event bus: what the gateway does, as a stream frontends can tap.
Events are plain dicts ``{"type", "seq", "ts", ...}``. Publishers never Events are plain dicts. Publishers never block; a subscriber that falls
block; a subscriber that falls behind loses its oldest events rather than behind loses its oldest events rather than stalling the publisher.
stalling the turn that produced them. ``/api/events`` serialises the
stream as SSE, the panel builds its subagent tree from ``stream`` events
carrying ``parent_tool_use_id``.
""" """
from __future__ import annotations from __future__ import annotations
+3 -7
View File
@@ -1,11 +1,7 @@
"""Unified streaming-event protocol shared by every backend adapter. """Streaming-event protocol shared by every backend adapter.
We piggyback on :mod:`anthropic.types` rather than reinvent the wire format: Aliases over :mod:`anthropic.types` plus ``build_*`` helpers that keep
``AnthropicMessagesFrontend`` ultimately serializes whatever a backend yields verbose constructors out of adapters.
straight to SSE via :py:meth:`pydantic.BaseModel.model_dump_json`, so events
must be valid Anthropic ``message_stream`` records. The aliases below give
the rest of the codebase one import path; the ``build_*`` helpers keep the
verbose constructors out of every adapter.
""" """
from __future__ import annotations from __future__ import annotations
+13 -17
View File
@@ -1,11 +1,7 @@
"""Collapse an Anthropic event stream into one ``Message``. """Collapse an Anthropic event stream into one ``Message``.
Extracted from ``AnthropicMessagesFrontend`` so the markdown frontend Mirrors the Anthropic SDK's own accumulator: folds content-block deltas
can run the same accumulation logic when it wants the finalized turn into finalized blocks, indexed by their ``content_block`` index.
rather than raw SSE chunks. Mirrors the Anthropic SDK's own accumulator:
walks events, builds block dicts indexed by their ``content_block``
index, folds text / thinking deltas in, buffers ``input_json_delta``
chunks until the block closes (then JSON-parses them once).
""" """
from __future__ import annotations from __future__ import annotations
@@ -42,9 +38,7 @@ __all__ = ["StreamAccumulator", "accumulate"]
class StreamAccumulator: class StreamAccumulator:
"""Folds a stream of events into one ``Message``, incrementally. """Folds a stream of events into one ``Message``, incrementally.
Use when you need to *both* forward events somewhere (SSE) *and* Call :meth:`feed` for each event, then :meth:`finalize` once.
keep a finalized ``Message`` for post-stream work (audit, logging
to disk). Call :meth:`feed` for each event, :meth:`finalize` once.
""" """
__slots__ = ( __slots__ = (
@@ -67,10 +61,11 @@ class StreamAccumulator:
self._stop_sequence: str | None = None self._stop_sequence: str | None = None
def feed(self, ev: MessageStreamEvent) -> None: def feed(self, ev: MessageStreamEvent) -> None:
# isinstance, not ``ev.type == "..."``: ty narrows on the """Fold one stream event into the in-progress message state.
# discriminator only via the class, and the raw event union
# carries its own discriminators (``Raw*Event``) the SDK Uses ``isinstance`` rather than ``ev.type == ...`` so ty narrows
# already promises. the event type from the class, not a string comparison.
"""
if isinstance(ev, RawMessageStartEvent): if isinstance(ev, RawMessageStartEvent):
self._message_id = ev.message.id self._message_id = ev.message.id
self._role = ev.message.role self._role = ev.message.role
@@ -109,6 +104,11 @@ class StreamAccumulator:
) )
def finalize(self, *, model: str) -> Message: def finalize(self, *, model: str) -> Message:
"""Build the finalized ``Message`` from accumulated block state.
``role`` is always ``"assistant"`` at the wire level; the cast
avoids a runtime check ty would otherwise require.
"""
content: list[Any] = [] content: list[Any] = []
for idx in sorted(self._blocks): for idx in sorted(self._blocks):
bd = self._blocks[idx] bd = self._blocks[idx]
@@ -120,10 +120,6 @@ class StreamAccumulator:
elif btype == "thinking": elif btype == "thinking":
content.append(ThinkingBlock.model_validate(bd)) content.append(ThinkingBlock.model_validate(bd))
# ``role`` is always ``"assistant"`` at the wire level — we
# initialised the field to that and only overwrite from a
# ``RawMessageStartEvent`` which itself carries the same literal.
# The cast keeps both type-checkers happy without a runtime check.
return Message( return Message(
id=self._message_id or "msg_unknown", id=self._message_id or "msg_unknown",
type="message", type="message",
+2 -10
View File
@@ -1,15 +1,7 @@
"""Admin console: serves the ``ui/`` SPA and signs the operator in. """Admin console: serves the ``ui/`` SPA and signs the operator in.
Everything the console shows comes from ``/api/*`` (``ApiFrontend``) Login/logout/session live under ``/admin/auth``; a process-lifetime
with a bearer. This app, mounted at ``/admin``, owns three JSON routes admin bearer (scope ``*``) is minted at startup and rotates on restart.
under ``/admin/auth`` - login (``ADMIN_USER`` / ``ADMIN_PASS`` from env,
session cookie signed with ``SESSION_SECRET``, 8 h), logout, and
``session``, which hands a signed-in browser the process-lifetime admin
bearer - and the static build under ``/admin/``.
The bearer is minted at startup, registered in the token store with
scope ``*`` and never persisted; a gateway restart rotates it, and the
SPA refetches ``session`` on a 401.
""" """
from __future__ import annotations from __future__ import annotations
+4 -15
View File
@@ -1,17 +1,8 @@
"""``POST /v1/messages`` frontend. """``POST /v1/messages`` frontend.
Exposes the gateway as an Anthropic-compatible Messages endpoint, so any Exposes the gateway as an Anthropic-compatible Messages endpoint, so any
client that already speaks Anthropic (Cursor, Cline, the official SDK, client that already speaks Anthropic can reach a configured agent by
``curl``) can hit a configured agent by passing its name as ``model``. passing its name as ``model``.
A Claude agent behind this endpoint is a ``deep`` conversation: the client
knows nothing about our ids, so the text fingerprint of the history it
sends is the ``(anthropic, fingerprint)`` binding of the conversation,
rebound after every turn to the fingerprint the next request will carry.
A history nobody has seen becomes a new conversation, materialized by the
home frontend of ``deep`` (the vault file), and every reply is published
on the bus so that file follows the chat. Other agents (Raycast) stay
stateless and are only archived through ``turn_log_handlers``.
""" """
from __future__ import annotations from __future__ import annotations
@@ -275,10 +266,8 @@ async def _sse(
) -> AsyncIterator[bytes]: ) -> AsyncIterator[bytes]:
r"""Serialize an event stream to SSE, then hand the assembled ``Message`` on. r"""Serialize an event stream to SSE, then hand the assembled ``Message`` on.
Each event becomes ``event: <type>\ndata: <json>\n\n`` - the shape Wire format is ``event: <type>\ndata: <json>\n\n``; mid-stream errors
the Anthropic SDK's SSE decoder expects. Errors mid-stream are become a synthetic ``error`` event instead of a hung connection.
swallowed into a synthetic ``error`` event so the client sees the
failure rather than a hung connection.
""" """
acc = StreamAccumulator() acc = StreamAccumulator()
try: try:
+1 -1
View File
@@ -1,4 +1,4 @@
"""``ApiFrontend`` - the conversations API and event stream (§3.9).""" """``ApiFrontend`` - the conversations API and event stream."""
from beaver_gateway.frontends.api.frontend import ApiFrontend from beaver_gateway.frontends.api.frontend import ApiFrontend
+2 -13
View File
@@ -1,18 +1,7 @@
"""``ApiFrontend`` - ``/api/*``: conversations, SSE, sessions, usage, limits (§3.9). """``ApiFrontend`` - ``/api/*``: conversations, SSE, sessions, usage, limits.
Bearer scope ``api``; token and audit management need ``admin``. Every Bearer scope ``api``; token and audit management need ``admin``. Every
write goes through ``core/conversations``; the frontend only shapes JSON. write goes through conversations/service.py; the frontend only shapes JSON.
``/api/events`` and ``/api/conversations/{id}/events`` replay the gateway
bus as SSE with the same keepalive the markdown frontend uses, so a proxy
never sees a silent socket.
Usage figures come from the ``usage`` table (one row per turn; API-price
``cost_usd`` and per-model ``model_usage`` are per-turn deltas of the SDK's
cumulative ``ResultMessage`` counters, see ``storage.append_usage``);
subscription quotas come from ``rate_limits``
(``RateLimitEvent``). The quota covers the whole subscription, so
``/api/limits`` puts the gateway's own spend for the window next to it
for calibration by eye.
""" """
from __future__ import annotations from __future__ import annotations
+17 -61
View File
@@ -1,13 +1,7 @@
"""Frontend ABC + the runtime context handed to ``configure``. """Frontend ABC and the runtime context handed to ``configure``.
A frontend is anything that routes inbound traffic into the gateway: an A frontend routes inbound traffic into the gateway (an HTTP mount, a
HTTP surface mounted under its ``path`` on the single gateway port poller, or both); ``GatewayRuntime`` carries the built state each needs.
(``frontends/root.py``), a poller (Telegram), or both. ``GatewayRuntime``
carries everything a frontend may need that isn't user-config: built
registries, per-agent backends, and the in-memory token store. The
user's ``/config/config.py`` defines a ``Gateway`` (lists); ``cli.main``
turns that into a ``GatewayRuntime`` and hands it to each frontend's
``configure``.
""" """
from __future__ import annotations from __future__ import annotations
@@ -36,20 +30,8 @@ if TYPE_CHECKING:
class GatewayRuntime: class GatewayRuntime:
"""Post-build state of the gateway, shared with every frontend. """Post-build state of the gateway, shared with every frontend.
Backends are keyed by **agent name**, not type one ``RaycastBackend`` Backends are keyed by agent name, not type: one backend instance may
instance can serve many ``RaycastAgent`` instances, but the lookup serve several agents, so the indirection lives at lookup time.
site (an inbound request with ``model=<agent.name>``) already has
the name in hand, so the indirection lives one step earlier.
``mcp_internal_urls`` is filled in Phase 2.1: one loopback URL per
declared ``McpServer`` so ``ClaudeSdkBackend``
can pass them to ``BackendOptions.mcp_servers`` without re-running
discovery.
``db`` (Phase 4.1) is the shared :class:`Database` handle. Phase 4.2
will switch ``TokenStore`` to read from it; Phase 4.3 admin/audit
write through it. Phase 4.1 only attaches it — existing frontends
ignore it.
""" """
agents: AgentRegistry agents: AgentRegistry
@@ -58,62 +40,36 @@ class GatewayRuntime:
token_store: TokenStore token_store: TokenStore
db: Database db: Database
mcp_internal_urls: Mapping[str, str] = field(default_factory=dict) mcp_internal_urls: Mapping[str, str] = field(default_factory=dict)
# Phase 4.3 — AdminFrontend reads creds + cookie-signing key from
# the runtime so the user's ``config.py`` doesn't have to know
# anything about env wiring. Defaulted to empty so existing tests /
# call sites that don't touch the admin path keep building; the
# admin frontend ``configure()`` itself rejects empty values.
admin_user: str = "" admin_user: str = ""
"""Operator login for the admin console, checked by ``AdminFrontend``."""
admin_pass: str = "" admin_pass: str = ""
session_secret: str = "" session_secret: str = ""
# The full sibling-frontends list, in declaration order. AdminFrontend
# uses it to advertise concrete bearer-endpoint URLs (host/port) on
# the dashboard so the operator can copy ready-to-use links / curl
# snippets. Other frontends ignore it.
frontends: Sequence[Frontend] = field(default_factory=tuple) frontends: Sequence[Frontend] = field(default_factory=tuple)
# Frontends that finish a turn (Anthropic Messages, Markdown) iterate """Every frontend in declaration order, for advertising their URLs."""
# this list and ``await`` each handler with a ``TurnRecord``. Handlers
# are appended during ``configure()`` by frontends that want a
# cross-frontend chat archive — currently the markdown frontend's
# ``log_all_chats`` mode. Handler exceptions are caught at the call
# site; they never block the user-visible response.
#
# The field is typed as ``list[Any]`` rather than the precise
# ``list[TurnLogHandler]`` because the alias lives under TYPE_CHECKING
# to keep ``anthropic.types`` out of the runtime import graph for
# this base module.
turn_log_handlers: list[TurnLogHandler] = field(default_factory=list) turn_log_handlers: list[TurnLogHandler] = field(default_factory=list)
# M1b: conversations service, event bus and the shared session pool. """Called with a ``TurnRecord`` after each turn; failures never reach the user."""
# ``Any`` for the same import-graph reason as above; ``None`` only in
# tests that build a runtime without them.
conversations: Any = None conversations: Any = None
bus: Any = None bus: Any = None
pool: Any = None pool: Any = None
scheduler: Any = None scheduler: Any = None
# External origin the reverse proxy puts in front of the gateway
# (``Gateway.public_url``); ``None`` means "derive from the request".
public_url: str | None = None public_url: str | None = None
"""``Gateway.public_url``; ``None`` derives the origin from the request."""
class Frontend(ABC): class Frontend(ABC):
"""Routes inbound traffic into the gateway. """Routes inbound traffic into the gateway.
HTTP frontends set ``path`` and return their ASGI app from ``app()``; HTTP frontends set ``path`` and return their ASGI app from ``app()``;
``cli`` mounts every such app under that path on the one gateway these are mounted under that path on the one gateway port. ``serve()``
port, so ``/anthropic/v1/messages`` reaches the Anthropic frontend's is for non-HTTP work (polling, vault mirrors) and defaults to nothing.
``/v1/messages``. ``serve()`` is for work outside HTTP - polling, ``landing`` marks the app that ``/`` redirects to.
vault mirrors - and defaults to nothing. ``landing`` marks the app
that ``/`` redirects to (the admin console).
A frontend that shows conversations declares ``name`` (the binding A frontend that shows conversations declares ``name`` (the binding
key) and ``kinds`` (which conversation kinds it shows); key) and ``kinds`` (which conversation kinds it shows). The first
``core/conversations`` refuses to bind a conversation to a frontend frontend whose ``materialize`` returns a binding is the *home* of
outside its declaration. The first frontend in declaration order that kind, used by ``spawn`` for new conversations; ``agent_for``
whose ``materialize`` returns a binding is the *home* of that kind: names the default agent for a kind. Stateless frontends (MCP, admin)
``spawn`` calls it so a new conversation gets a window (a vault file, leave these at their defaults.
a Telegram topic). ``agent_for`` names the default agent for a kind
so callers may omit ``agent``. Stateless frontends (MCP, admin) keep
the defaults and stay outside the routing.
""" """
name: str = "" name: str = ""
+6 -12
View File
@@ -1,8 +1,7 @@
"""Shared bearer-token verification for HTTP frontends. """Shared bearer-token verification for HTTP frontends.
Extracted from ``AnthropicMessagesFrontend`` so the markdown frontend Reused by every bearer-protected frontend instead of duplicating the
(and any future bearer-protected frontend) can reuse one canonical header-parsing logic.
verifier instead of copy-pasting the header-parsing dance.
""" """
from __future__ import annotations from __future__ import annotations
@@ -23,14 +22,11 @@ __all__ = ["require_token"]
async def require_token( async def require_token(
request: Request, runtime: GatewayRuntime, *, scope: str request: Request, runtime: GatewayRuntime, *, scope: str
) -> str: ) -> str:
"""Verify the request's bearer + scope, return the token's audit name. """Verify the request's bearer token and scope; return its audit name.
Accepts ``X-Api-Key: <token>`` (Anthropic SDK / LibreChat), Checks ``X-Api-Key``, then ``Authorization: Bearer``, then ``?token=``.
``Authorization: Bearer <token>`` (curl, Cursor) and, when neither 401 on a missing/unknown token, 403 if the token's scope doesn't cover
header is present, ``?token=<token>`` (Komodo alerters). 401 on missing / ``scope``. Bootstrap tokens carry ``"*"`` and pass every scope check.
unknown token; 403 on a known token whose scope doesn't cover
``scope``. Bootstrap tokens implicitly carry ``"*"`` and pass every
scope check.
""" """
api_key = request.headers.get("x-api-key") api_key = request.headers.get("x-api-key")
authorization = request.headers.get("authorization") authorization = request.headers.get("authorization")
@@ -39,8 +35,6 @@ async def require_token(
elif authorization: elif authorization:
identity = await runtime.token_store.verify_bearer(authorization) identity = await runtime.token_store.verify_bearer(authorization)
else: else:
# Webhook senders that cannot set headers (Komodo alerters) put the
# token in the query string; the URL is not logged with it.
qs_token = request.query_params.get("token") qs_token = request.query_params.get("token")
identity = await runtime.token_store.verify(qs_token) if qs_token else None identity = await runtime.token_store.verify(qs_token) if qs_token else None
if identity is None: if identity is None:
@@ -1,13 +1,7 @@
"""Markdown frontend — turn-by-turn chat archive backed by ``.md`` files. """Markdown frontend — chat archive backed by ``.md`` files in an Obsidian vault.
The user maintains chats as plain markdown files in an Obsidian vault. An Obsidian plugin POSTs ``{filename, content?}`` to ``/chat``; the
A plugin in Obsidian POSTs ``{filename, content?}`` to ``/chat`` and the frontend runs the agent on a trailing user turn and appends the reply.
frontend parses the file, finds the last turn, and runs the agent if
the last turn is ``user``. The full response is appended back to the
file as an ``### Assistant:`` turn. With ``log_all_chats=True`` the
frontend also subscribes to every other frontend's turns and writes
them into ``{vault_path}/{logged_subdir}/`` so the vault is the single
chronological archive of all conversations.
""" """
from beaver_gateway.frontends.markdown.frontend import MarkdownFrontend from beaver_gateway.frontends.markdown.frontend import MarkdownFrontend
@@ -1,17 +1,7 @@
"""Cross-frontend chat logger. """Cross-frontend chat logger.
When ``MarkdownFrontend(log_all_chats=True)`` is configured, every turn Mirrors turns completed by other frontends into the vault as ``.md``
completed by any other frontend (currently the Anthropic Messages files, matching a conversation's continuation by content-hash fingerprint.
frontend) is mirrored into the vault as a ``.md`` file. Subsequent
turns of the same conversation append to the same file — matched by a
content-hash fingerprint stored in YAML frontmatter.
The fingerprint hashes the message history *before* the new assistant
reply. So the next request's input history (which now includes the
prior assistant reply) hashes to the value we just persisted —
``hash(prev_input + [assistant_reply])`` — and the lookup hits the
same file. New conversations (no prior fingerprint match) get a fresh
file under ``{vault_path}/{logged_subdir}/{agent_name}/``.
""" """
from __future__ import annotations from __future__ import annotations
@@ -53,12 +43,8 @@ vault. ``None`` keeps ``{vault}/{logged_subdir}/{agent}/{YYYY-MM-DD}_{slug}.md``
def fingerprint_messages(messages: Iterable[MessageParam]) -> str: def fingerprint_messages(messages: Iterable[MessageParam]) -> str:
"""Stable, short hex hash of a conversation prefix. """Stable, short hex hash of a conversation prefix.
Built from ``(role, normalized_content)`` pairs only — so the Hashes ``(role, text-only content)`` pairs so differently-shaped message
Markdown frontend's parser-shaped messages (text-only) and the histories that carry the same text still fingerprint identically.
Anthropic frontend's raw ``messages`` payload (which may also be
string-only at v1) hash compatibly when they represent the same
conversation. Tool blocks / images would diverge, but those aren't
in the v1 ingest path.
""" """
h = hashlib.sha1(usedforsecurity=False) h = hashlib.sha1(usedforsecurity=False)
for msg in messages: for msg in messages:
@@ -85,12 +71,8 @@ def fingerprint_messages(messages: Iterable[MessageParam]) -> str:
class CrossFrontendLogger: class CrossFrontendLogger:
"""Maintains the fingerprint→file map and writes turns to disk. """Maintains the fingerprint→file map and writes turns to disk.
The map is in-process; on startup ``warm_index`` rebuilds it from The map is in-process, rebuilt by ``warm_index`` on startup; all disk
YAML frontmatter of every file under ``logged_subdir``. A miss writes funnel through one lock to sidestep races.
creates a new file, a hit appends to the existing one. All disk
work funnels through one ``asyncio.Lock`` because the writes are
cheap and serializing them sidesteps a class of races we don't need
to think about.
""" """
def __init__( def __init__(
@@ -105,20 +87,13 @@ class CrossFrontendLogger:
self._index: dict[str, Path] = {} self._index: dict[str, Path] = {}
self._lock = asyncio.Lock() self._lock = asyncio.Lock()
self._chat_path_fn = chat_path self._chat_path_fn = chat_path
# When the user supplies a custom path function, files can land
# anywhere in the vault — so we have to scan the whole vault on
# startup to rebuild the fingerprint→path map. With the default
# layout we can bound the scan to ``_logs/``.
self._scan_root = vault_path if chat_path is not None else self._root self._scan_root = vault_path if chat_path is not None else self._root
def warm_index(self) -> None: def warm_index(self) -> None:
"""Scan logged files synchronously, populating the fingerprint map. """Scan logged files synchronously, populating the fingerprint map.
Called from ``MarkdownFrontend.configure`` so the map is ready Called before any cross-frontend turn arrives; a custom ``chat_path``
before any cross-frontend turn arrives. ``frontmatter.load`` forces a full-vault walk instead of scanning just ``logged_subdir``.
reads only enough of the file to parse the YAML head, so the
scan is cheap even on large vaults — but a custom ``log_path``
forces a full-vault walk; mention that in the constructor doc.
""" """
if not self._scan_root.exists(): if not self._scan_root.exists():
return return
@@ -140,29 +115,19 @@ class CrossFrontendLogger:
async def handle(self, record: TurnRecord) -> None: async def handle(self, record: TurnRecord) -> None:
"""Append or create a logged file for ``record``. """Append or create a logged file for ``record``.
Records that the markdown frontend itself produced Skips ``source == "markdown"`` records (already on disk); matches
(``source=="markdown"``) are skipped — those already live in the the target file by fingerprinting ``input_messages`` sans the new turn.
user's hand-written file and shouldn't be duplicated into the
``_logs`` shadow tree.
""" """
if record.source == "markdown": if record.source == "markdown":
return return
async with self._lock: async with self._lock:
# ``input_messages`` is the *full* history sent to the backend
# (last entry is the new user turn). Match against the prefix
# that excludes the new user turn — that's what the previous
# write stored as its fingerprint. Empty prefix is the
# well-known "brand new chat" sentinel.
prefix = record.input_messages[:-1] prefix = record.input_messages[:-1]
prev_fp = fingerprint_messages(prefix) if prefix else None prev_fp = fingerprint_messages(prefix) if prefix else None
target = self._index.get(prev_fp) if prev_fp else None target = self._index.get(prev_fp) if prev_fp else None
if target is None: if target is None:
target = self._new_file_path(record) target = self._new_file_path(record)
# Build the full history including the assistant reply; the
# new fingerprint matches *that* prefix, so the next user
# turn (history grows by one user msg) will hit this file.
assistant_msg: MessageParam = { assistant_msg: MessageParam = {
"role": "assistant", "role": "assistant",
"content": _flatten_text(record.output_message), "content": _flatten_text(record.output_message),
@@ -174,9 +139,6 @@ class CrossFrontendLogger:
existing = target.read_text(encoding="utf-8") existing = target.read_text(encoding="utf-8")
parsed = frontmatter.loads(existing) parsed = frontmatter.loads(existing)
body = strip_trailing_user_scaffold(parsed.content) body = strip_trailing_user_scaffold(parsed.content)
# We append only the *new* user turn (the last one in
# input_messages, since prior turns are already on disk)
# plus the assistant reply.
new_user = record.input_messages[-1] new_user = record.input_messages[-1]
new_block = renderer.render_user_param(new_user) new_block = renderer.render_user_param(new_user)
new_block = renderer.append_to_body( new_block = renderer.append_to_body(
@@ -185,7 +147,6 @@ class CrossFrontendLogger:
new_body = renderer.append_to_body(body, new_block) new_body = renderer.append_to_body(body, new_block)
metadata = dict(parsed.metadata) metadata = dict(parsed.metadata)
else: else:
# Materialize the whole conversation from scratch.
new_body = _render_full_history( new_body = _render_full_history(
record.input_messages, record.output_message record.input_messages, record.output_message
) )
@@ -197,21 +158,15 @@ class CrossFrontendLogger:
metadata["fingerprint"] = new_fp metadata["fingerprint"] = new_fp
metadata["source"] = record.source metadata["source"] = record.source
self._write(target, metadata, new_body) self._write(target, metadata, new_body)
# Maintain the index: drop the old fp (it's stale once we
# write the new turn), add the new one.
if prev_fp: if prev_fp:
self._index.pop(prev_fp, None) self._index.pop(prev_fp, None)
self._index[new_fp] = target self._index[new_fp] = target
# ---- internals -----------------------------------------------------
def _new_file_path(self, record: TurnRecord) -> Path: def _new_file_path(self, record: TurnRecord) -> Path:
"""Pick a fresh filename for a brand-new conversation. """Pick a fresh filename for a brand-new conversation.
With a user-supplied ``chat_path`` we delegate to it (joining a Delegates to ``chat_path`` if set; otherwise
relative result with the vault root). Without one, we fall back ``{logged_subdir}/{agent}/{date}_{hex8}.md``.
to ``{logged_subdir}/{agent}/{date}_{hex8}.md`` and ensure the
``.md`` suffix in case the user picks a non-md extension by hand.
""" """
if self._chat_path_fn is not None: if self._chat_path_fn is not None:
result = self._chat_path_fn( result = self._chat_path_fn(
@@ -224,8 +179,6 @@ class CrossFrontendLogger:
result.parent.mkdir(parents=True, exist_ok=True) result.parent.mkdir(parents=True, exist_ok=True)
return result return result
day = datetime.now(UTC).strftime("%Y-%m-%d") day = datetime.now(UTC).strftime("%Y-%m-%d")
# Short hex from the input hash so two same-day chats sort
# stably and don't collide.
salt = fingerprint_messages(record.input_messages)[:8] salt = fingerprint_messages(record.input_messages)[:8]
agent_dir = self._root / record.agent_name agent_dir = self._root / record.agent_name
agent_dir.mkdir(parents=True, exist_ok=True) agent_dir.mkdir(parents=True, exist_ok=True)
@@ -238,9 +191,6 @@ class CrossFrontendLogger:
text = frontmatter.dumps(post) + "\n" text = frontmatter.dumps(post) + "\n"
else: else:
text = body if body.endswith("\n") else body + "\n" text = body if body.endswith("\n") else body + "\n"
# Sync write inside the lock — keeps the implementation tiny;
# individual logged turns are small enough that the blocking
# write doesn't matter at human conversation rates.
path.write_text(text, encoding="utf-8") path.write_text(text, encoding="utf-8")
@@ -277,18 +227,13 @@ def _render_full_history(messages: list[MessageParam], assistant: Any) -> str:
def strip_trailing_user_scaffold(body: str) -> str: def strip_trailing_user_scaffold(body: str) -> str:
"""Drop a trailing empty ``### User:`` block if present. """Drop a trailing empty ``### User:`` block if present.
Cross-frontend turns aren't typed into the file by the human — they Avoids leaving two ``### User:`` headers in a row when appending a
arrive whole from another frontend. If we leave the previous run's turn that wasn't typed into the file by hand.
scaffold in place, we'd write the new user turn right after an
empty marker (visual noise, two ``### User:`` headers in a row).
Trim it and let the append flow add a fresh scaffold at the end.
""" """
stripped = body.rstrip() stripped = body.rstrip()
marker = "### User:" marker = "### User:"
if not stripped.endswith(marker): if not stripped.endswith(marker):
return body return body
# Walk back: the scaffold is the marker preceded by either start-of-file
# or an HR/blank line. Find the last newline before the marker, cut.
head = stripped[: -len(marker)].rstrip() head = stripped[: -len(marker)].rstrip()
if head.endswith("---"): if head.endswith("---"):
head = head[: -len("---")].rstrip() head = head[: -len("---")].rstrip()
+35 -131
View File
@@ -1,34 +1,7 @@
"""``MarkdownFrontend`` — chat-via-markdown-files frontend. """``MarkdownFrontend`` — chat-via-markdown-files frontend.
Wires: ``POST /chat`` (and SSE ``/chat/stream``) parses the vault file, runs
the agent on the trailing user turn, and appends the reply.
* ``POST /chat {filename, content?, agent?}`` — bearer-authenticated
trigger. The plugin in Obsidian fires this after the user edits a
``.md`` and the file gets synced (or with ``content`` to short-circuit
the sync delay). We parse the file, check the last turn — assistant
→ no-op, user → run the agent and append.
* ``GET /healthz`` — liveness.
Concurrency model: an in-memory ``set[Path]`` of files currently in
flight. Two concurrent requests for the same file → the second gets
409. The set is single-process (one gateway instance) — that's by
design; the markdown frontend is the only writer in its vault from
the gateway side. The turn itself runs through ``core/conversations``
(one turn per conversation, ``running_turn`` in the DB), so a message
posted to the same conversation via ``/api`` waits its turn.
A chat file is a ``deep`` conversation bound as
``(markdown, <vault-relative path>)``; frontmatter carries only ``agent``
and ``conversation_id`` (§3.10), tool calls are never rendered. The
frontend is the home of ``deep``: ``materialize`` gives a conversation
spawned elsewhere its file, and :class:`.mirror.ChatMirror` keeps that
file in step with replies produced outside ``/chat``.
Cross-frontend logging: when ``log_all_chats=True``, ``configure()``
registers a handler on ``runtime.turn_log_handlers`` so every other
frontend's completed turns also land in the vault. The handler logic
lives in :mod:`.crossfront` so this module stays focused on the HTTP
shape.
""" """
from __future__ import annotations from __future__ import annotations
@@ -85,23 +58,17 @@ _log = logging.getLogger("beaver_gateway.frontends.markdown")
__all__ = ["MarkdownFrontend"] __all__ = ["MarkdownFrontend"]
# How often we re-render the assistant turn into the .md file while the
# backend stream is still open. Trades responsiveness (faster updates to
# Obsidian sync / Raycast tailers) against write amplification. Each
# ``RawContentBlockStopEvent`` also forces a flush regardless of the
# timer, so block boundaries always land in the file.
_STREAM_FLUSH_DEBOUNCE = 0.4 _STREAM_FLUSH_DEBOUNCE = 0.4
# Debounce for the SSE ``/chat/stream`` path. Network IO is cheaper than
# atomic file rewrites, so we send updates more frequently — the client
# wants the lowest possible latency and we control the renderer on the
# other end (the Obsidian plugin splices deltas into the editor, no
# disk round-trip).
_SSE_FLUSH_DEBOUNCE = 0.1 _SSE_FLUSH_DEBOUNCE = 0.1
class MarkdownFrontend(Frontend): class MarkdownFrontend(Frontend):
"""FastAPI app behind ``POST /chat`` driven by Obsidian-vault files.""" """FastAPI app behind ``POST /chat`` driven by Obsidian-vault files.
``_busy`` tracks in-flight files; check-and-add must stay atomic (no
``await`` between them) so concurrent requests reliably lose to 409.
"""
name = FRONTEND name = FRONTEND
kinds = ("deep",) kinds = ("deep",)
@@ -116,6 +83,11 @@ class MarkdownFrontend(Frontend):
logged_subdir: str = "_logs", logged_subdir: str = "_logs",
chat_path: Callable[[str, str, Path], Path] | None = None, chat_path: Callable[[str, str, Path], Path] | None = None,
) -> None: ) -> None:
"""Configure the vault-backed frontend.
``chat_path``, if given, overrides where new chat files are created;
``logged_subdir`` holds cross-frontend logs when ``log_all_chats`` is set.
"""
self.vault_path = Path(vault_path).expanduser().resolve() self.vault_path = Path(vault_path).expanduser().resolve()
self.default_agent = default_agent self.default_agent = default_agent
self.log_all_chats = log_all_chats self.log_all_chats = log_all_chats
@@ -123,10 +95,6 @@ class MarkdownFrontend(Frontend):
self.chat_path = chat_path self.chat_path = chat_path
self._runtime: GatewayRuntime | None = None self._runtime: GatewayRuntime | None = None
self._app: FastAPI | None = None self._app: FastAPI | None = None
# Files currently being processed by an in-flight ``POST /chat``.
# Checked-and-added atomically in the request handler (no
# ``await`` between the check and the insert) so a concurrent
# request reliably loses the race to 409.
self._busy: set[Path] = set() self._busy: set[Path] = set()
self._crossfront: CrossFrontendLogger | None = None self._crossfront: CrossFrontendLogger | None = None
self._mirror: ChatMirror | None = None self._mirror: ChatMirror | None = None
@@ -149,9 +117,6 @@ class MarkdownFrontend(Frontend):
logged_subdir=self.logged_subdir, logged_subdir=self.logged_subdir,
chat_path=self.chat_path, chat_path=self.chat_path,
) )
# Scan the existing logged files synchronously here so the
# fingerprint→path map is populated before the first
# cross-frontend turn arrives. Cheap: frontmatter-only read.
self._crossfront.warm_index() self._crossfront.warm_index()
runtime.turn_log_handlers.append(self._crossfront.handle) runtime.turn_log_handlers.append(self._crossfront.handle)
self._app = self._build_app(runtime) self._app = self._build_app(runtime)
@@ -175,16 +140,10 @@ class MarkdownFrontend(Frontend):
async def serve(self) -> None: async def serve(self) -> None:
await self.mirror.run() await self.mirror.run()
# ---- app builder ---------------------------------------------------
def _build_app(self, runtime: GatewayRuntime) -> FastAPI: def _build_app(self, runtime: GatewayRuntime) -> FastAPI:
"""CORS is wide open here since auth is bearer-token, not cookie-based."""
app = FastAPI(title="beaver-gateway / Markdown") app = FastAPI(title="beaver-gateway / Markdown")
# ``/chat/stream`` is consumed via ``fetch`` from the Obsidian
# plugin (``requestUrl`` can't read a body incrementally), and
# ``fetch`` is subject to CORS. Auth is bearer-token so we don't
# need credentialed mode; allow any origin and the standard
# methods/headers. The other endpoints are happy to ride along.
app.add_middleware( app.add_middleware(
CORSMiddleware, CORSMiddleware,
allow_origins=["*"], allow_origins=["*"],
@@ -226,8 +185,6 @@ class MarkdownFrontend(Frontend):
file_path = self._resolve_path(filename) file_path = self._resolve_path(filename)
# Atomic check-and-claim: both ops run between awaits, so a
# second request can't slip into the same file slot.
if file_path in self._busy: if file_path in self._busy:
return JSONResponse( return JSONResponse(
status_code=status.HTTP_409_CONFLICT, status_code=status.HTTP_409_CONFLICT,
@@ -248,12 +205,10 @@ class MarkdownFrontend(Frontend):
@app.post("/chat/stream") @app.post("/chat/stream")
async def chat_stream(request: Request) -> Any: async def chat_stream(request: Request) -> Any:
# Same contract as ``/chat`` (bearer auth, identical body), """SSE variant of ``/chat``: ``delta`` events, one disk write at the end.
# but the response is ``text/event-stream`` and intermediate
# rendered states are pushed as ``delta`` events. The The 409-conflict response is still plain JSON — the stream hasn't started.
# gateway-side disk write only happens once, at end of turn, """
# so streaming consumers (Obsidian plugin) and Obsidian Sync
# don't fight over the same file mid-stream.
token_name = await require_token(request, runtime, scope="messages") token_name = await require_token(request, runtime, scope="messages")
try: try:
body = await request.json() body = await request.json()
@@ -276,8 +231,6 @@ class MarkdownFrontend(Frontend):
file_path = self._resolve_path(filename) file_path = self._resolve_path(filename)
# 409 path stays JSON — the stream hasn't started yet, so
# the caller can read it the same way as on ``/chat``.
if file_path in self._busy: if file_path in self._busy:
return JSONResponse( return JSONResponse(
status_code=status.HTTP_409_CONFLICT, status_code=status.HTTP_409_CONFLICT,
@@ -305,8 +258,6 @@ class MarkdownFrontend(Frontend):
return app return app
# ---- dispatch ------------------------------------------------------
async def _handle_chat( async def _handle_chat(
self, self,
*, *,
@@ -317,6 +268,7 @@ class MarkdownFrontend(Frontend):
content_override: Any, content_override: Any,
agent_override: str | None, agent_override: str | None,
) -> Any: ) -> Any:
"""Non-streaming ``/chat`` handler: parse, align, run, persist."""
write_disk = content_override is None write_disk = content_override is None
if isinstance(content_override, str): if isinstance(content_override, str):
file_text = content_override file_text = content_override
@@ -340,8 +292,6 @@ class MarkdownFrontend(Frontend):
"or configure `default_agent`", "or configure `default_agent`",
) )
# When the parser produced no messages (file empty / only
# frontmatter), there's nothing to dispatch.
if not parsed.messages: if not parsed.messages:
return { return {
"status": "nothing_to_do", "status": "nothing_to_do",
@@ -384,9 +334,6 @@ class MarkdownFrontend(Frontend):
msgs=len(parsed.messages), msgs=len(parsed.messages),
) )
# Resolve / mint the conversation row, align incoming against
# stored history, and feed the aligned messages to the backend
# - see ``core/conversation_store.py`` for the full rationale.
conv, conv_external_id, stored_msgs = await self._resolve_conversation( conv, conv_external_id, stored_msgs = await self._resolve_conversation(
runtime=runtime, runtime=runtime,
metadata=parsed.metadata, metadata=parsed.metadata,
@@ -437,9 +384,6 @@ class MarkdownFrontend(Frontend):
message=message, message=message,
) )
# Broadcast our own turn so other handlers (none today, but the
# symmetry is worth keeping) see what happened. ``source`` marks
# the origin so ``CrossFrontendLogger`` can skip its own files.
record = TurnRecord( record = TurnRecord(
agent_name=agent.name, agent_name=agent.name,
input_messages=list(parsed.messages), input_messages=list(parsed.messages),
@@ -460,9 +404,7 @@ class MarkdownFrontend(Frontend):
"new_content": new_content, "new_content": new_content,
} }
# ---- streaming dispatch (SSE) -------------------------------------- async def _handle_chat_streaming( # noqa: PLR0915
async def _handle_chat_streaming( # noqa: PLR0915 — mirrors _handle_chat, splitting only doubles read cost
self, self,
*, *,
runtime: GatewayRuntime, runtime: GatewayRuntime,
@@ -474,21 +416,9 @@ class MarkdownFrontend(Frontend):
) -> AsyncIterator[bytes]: ) -> AsyncIterator[bytes]:
"""SSE counterpart of :meth:`_handle_chat`. """SSE counterpart of :meth:`_handle_chat`.
Mirrors the same pipeline (resolve file → parse → resolve agent → Errors surface as ``error`` frames (HTTP is already 200 by then); disk
run backend → persist), but emits ``event: delta`` frames as the is only written once at the end, avoiding a write race with Obsidian Sync.
rendered turn grows and a single terminal ``event: done`` /
``event: error``. Errors that ``_handle_chat`` would surface as
``HTTPException`` go out as ``error`` frames here (the HTTP
envelope is already 200 by the time the stream starts).
Intermediate disk writes are deliberately skipped — only the
post-stream :meth:`_write_assistant_reply` lands on disk, so the
gateway-side vault and the plugin-side editor are the only
writers in their respective halves of Obsidian Sync. Final
content is identical on both sides, so Sync no-ops.
""" """
# With ``content`` the plugin is the only writer of the file
# (§3.10): the gateway never touches disk in that case.
write_disk = content_override is None write_disk = content_override is None
if isinstance(content_override, str): if isinstance(content_override, str):
file_text = content_override file_text = content_override
@@ -653,18 +583,12 @@ class MarkdownFrontend(Frontend):
or (now - last_flush) >= _SSE_FLUSH_DEBOUNCE or (now - last_flush) >= _SSE_FLUSH_DEBOUNCE
): ):
payload = snapshot() payload = snapshot()
# Skip duplicate snapshots — e.g. tool_use blocks
# render to the same prefix as before they closed
# (we don't surface the tool-call args in markdown).
if payload is not None and payload != last_payload: if payload is not None and payload != last_payload:
yield sse_pack("delta", {"new_content": payload}) yield sse_pack("delta", {"new_content": payload})
last_payload = payload last_payload = payload
last_flush = now last_flush = now
except Exception as exc: # noqa: BLE001 — wire any backend failure as an SSE error frame except Exception as exc: # noqa: BLE001
_log.exception("backend failed for %s", filename) _log.exception("backend failed for %s", filename)
# Mirror the legacy path: write the last partial + an error
# callout to disk so other consumers (logs, file watchers)
# see what arrived. The client gets a clean SSE ``error``.
partial = acc.finalize(model=model) partial = acc.finalize(model=model)
new_body = parsed.body new_body = parsed.body
if partial.content: if partial.content:
@@ -728,8 +652,6 @@ class MarkdownFrontend(Frontend):
}, },
) )
# ---- helpers -------------------------------------------------------
async def _stream_to_file( async def _stream_to_file(
self, self,
*, *,
@@ -742,17 +664,8 @@ class MarkdownFrontend(Frontend):
) -> Any: ) -> Any:
"""Drain ``events`` into a ``Message``, flushing partials to disk. """Drain ``events`` into a ``Message``, flushing partials to disk.
Flushes happen on each ``RawContentBlockStopEvent`` (natural Flushes on each block boundary and on a debounce timer; on backend
block boundary, content is markdown-consistent) and on the failure it still flushes a partial + error callout, then re-raises.
``_STREAM_FLUSH_DEBOUNCE`` timer between events. The partial
write keeps the as-parsed frontmatter; the post-stream final
write in ``_write_assistant_reply`` is what stamps the refreshed
fingerprint / agent / conversation_id.
On backend exception we still flush the last partial and append
an error callout, so the human sees both what arrived and why it
stopped. The exception propagates so ``_handle_chat`` can map it
to a 500.
""" """
acc = StreamAccumulator() acc = StreamAccumulator()
@@ -829,11 +742,8 @@ class MarkdownFrontend(Frontend):
) -> tuple[Conversation, str, list[dict[str, Any]]]: ) -> tuple[Conversation, str, list[dict[str, Any]]]:
"""Resolve the ``deep`` conversation for this file + its stored messages. """Resolve the ``deep`` conversation for this file + its stored messages.
Frontmatter ``conversation_id`` wins; a file that lost it is found Precedence: frontmatter ``conversation_id`` > existing path binding >
by its visible ``(markdown, path)`` binding, then by adopting the an adopted unbound conversation with the same first prompt > a new one.
one unbound pre-SDK conversation that starts with the same prompt;
otherwise a new conversation is created. The binding follows the
file: a moved chat re-binds to its new path on the next turn.
""" """
conversations = runtime.conversations conversations = runtime.conversations
rel = file_path.relative_to(self.vault_path).as_posix() rel = file_path.relative_to(self.vault_path).as_posix()
@@ -882,9 +792,8 @@ class MarkdownFrontend(Frontend):
) -> None: ) -> None:
"""Stamp the DB with the post-turn canonical Anthropic-shape history. """Stamp the DB with the post-turn canonical Anthropic-shape history.
Combines the matched/spliced prior state, the new user prompt, Combines prior state + new user prompt + the backend's synthesized
and the synthesized assistant/tool cycle from the backend (or a cycle, falling back to text-only if ``capture`` is empty.
text-only fallback for backends that left ``capture`` empty).
""" """
new_user_msg = {"role": "user", "content": new_user_text} new_user_msg = {"role": "user", "content": new_user_text}
synthesized = capture.synthesized_messages or _fallback_synthesized(message) synthesized = capture.synthesized_messages or _fallback_synthesized(message)
@@ -906,12 +815,11 @@ class MarkdownFrontend(Frontend):
) )
def _resolve_path(self, filename: str) -> Path: def _resolve_path(self, filename: str) -> Path:
"""Resolve ``filename`` under the vault; reject escapes.""" """Resolve ``filename`` under the vault; reject escapes.
# ``filename`` may be relative or absolute; we always anchor
# under ``vault_path`` so absolute paths from outside the vault Leading slashes are stripped first — ``Path.__truediv__`` would
# don't sneak through. ``Path("/foo/bar")`` combined with a otherwise discard ``vault_path`` for an absolute ``filename``.
# vault path keeps the absolute side; we strip leading slashes """
# to coerce the rooted form into a relative path before joining.
rel = filename.lstrip("/") rel = filename.lstrip("/")
if not rel.endswith(".md"): if not rel.endswith(".md"):
rel = rel + ".md" rel = rel + ".md"
@@ -928,12 +836,8 @@ class MarkdownFrontend(Frontend):
def _fallback_synthesized(message: Any) -> list[dict[str, Any]]: def _fallback_synthesized(message: Any) -> list[dict[str, Any]]:
"""Build a single-assistant ``synthesized_messages`` list from a raw ``Message``. """Build a single-assistant ``synthesized_messages`` list from a raw ``Message``.
For backends that don't populate a :class:`TurnCapture` (anthropic Fallback for backends that don't populate :class:`TurnCapture` (anthropic
HTTP, raycast, …) we don't have access to per-tool-cycle HTTP, raycast, …); the reply lands as one canonical-block message.
granularity, so the assistant reply lands in the DB as one
canonical-block message. Tool memory across cache misses would
degrade in that case, but those backends don't have the cache-miss
re-seed problem to begin with — they manage history client-side.
""" """
content: list[dict[str, Any]] = [] content: list[dict[str, Any]] = []
for block in getattr(message, "content", ()): for block in getattr(message, "content", ()):
+34 -180
View File
@@ -1,41 +1,7 @@
"""Stateful conversation history for the markdown frontend. """Stored conversation history for the markdown frontend.
The gateway used to be stateless about identity: claude-code-api's Aligns the file the user is editing against the DB-stored Anthropic-shape
in-memory session pool was keyed by a fingerprint of the messages the message history, splicing text edits or forking on structural changes.
gateway forwarded, and on a fingerprint miss the same fingerprint was
used to seed a fresh PTY's JSONL transcript. That worked as long as
the frontend could round-trip the *exact* content blocks the live
session had observed. The markdown frontend can't — the parser strips
``[!tool]-`` callouts because the human is allowed to edit the prose,
and the rendered tool callouts don't carry the canonical ``tool_use``
block fields anyway. So a continuation hit was *only* reliable for
turns that never used a tool; once tools entered the picture, every
subsequent turn missed the cache and reseeded from a tool-less
transcript, leading to "assistant doesn't remember the tool calls it
just made."
This module makes the gateway stateful for the markdown frontend (and
any other frontend that wants in). The DB stores the full
Anthropic-shape message list — text blocks, ``tool_use`` blocks,
``tool_result`` blocks, thinking signatures — exactly as
claude-code-api would have seen on the wire. Before each turn we
align the file the user is editing against the stored history:
* If the user just appended a new user turn at the bottom, we feed
the backend our stored-plus-new history and the fingerprint hits.
* If the user edited the *text* inside an assistant turn but left the
tool callouts alone, we splice the new text into the stored
``tool_use`` blocks and feed *that* — the fingerprint misses (text
differs), claude-code-api reseeds with a full transcript (tools and
all), the new live session has memory of the prior tool calls.
* If the user changed the *structure* (added/removed/reordered a tool
callout, edited an old user turn, etc.) we fork: take stored history
up to the divergence, take incoming text-only past the divergence.
The fingerprint misses; claude-code-api reseeds with a clean
truncated history; downstream turns continue from there.
"Divergence point" is found by walking the file's turns and the
stored display turns in lockstep. See :func:`diff_and_fork`.
""" """
from __future__ import annotations from __future__ import annotations
@@ -67,21 +33,12 @@ __all__ = [
] ]
# ---- types --------------------------------------------------------------
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
class ForkOutcome: class ForkOutcome:
"""Result of aligning the incoming file against stored history. """Result of aligning the incoming file against stored history.
``messages`` is what the gateway feeds to the backend (already ``messages`` is the backend input; ``persist_messages`` is history to
includes the new user prompt at the tail). ``persist_messages`` store before the new reply; ``divergence_index`` is where they diverged.
is the canonical conversation state the gateway should hold in
the DB *up to but not including* the new assistant reply — the
caller appends the synthesized turn from the backend onto this
and writes the result back. ``divergence_index`` is the
display-turn index at which incoming first disagreed with stored
(``None`` if everything matched; the new tail is appended cleanly).
""" """
messages: list[MessageParam] messages: list[MessageParam]
@@ -95,9 +52,6 @@ class ForkOutcome:
return self.divergence_index is None and not self.edited return self.divergence_index is None and not self.edited
# ---- public store API ---------------------------------------------------
async def load_conversation( async def load_conversation(
session: AsyncSession, *, frontend: str, external_id: str session: AsyncSession, *, frontend: str, external_id: str
) -> Conversation | None: ) -> Conversation | None:
@@ -115,9 +69,7 @@ async def mint_conversation(
) -> Conversation: ) -> Conversation:
"""Create a fresh conversation row with a new uuid for external_id. """Create a fresh conversation row with a new uuid for external_id.
Caller is responsible for persisting the returned ``external_id`` on Caller must persist the returned ``external_id`` so future requests can find it.
the frontend side (frontmatter, response header, …) so future
requests can find this conversation again.
""" """
row = Conversation( row = Conversation(
frontend=frontend, external_id=str(uuid.uuid4()), agent_name=agent_name frontend=frontend, external_id=str(uuid.uuid4()), agent_name=agent_name
@@ -144,9 +96,8 @@ async def load_messages(
) -> list[dict[str, Any]]: ) -> list[dict[str, Any]]:
"""Return stored messages ordered by ``seq`` ascending. """Return stored messages ordered by ``seq`` ascending.
Each entry is a canonical Anthropic ``MessageParam`` dict — ``role`` Each entry is a canonical Anthropic ``MessageParam`` dict, the same
plus ``content`` (string or list of block dicts). The same shape shape fed to the backend on continuation.
we feed to the backend on continuation.
""" """
stmt = ( stmt = (
select(ConversationMessage) select(ConversationMessage)
@@ -164,10 +115,8 @@ async def load_messages(
def _sanitize_content(content: Any) -> Any: def _sanitize_content(content: Any) -> Any:
"""Strip wire-illegal fields from stored Anthropic content blocks. """Strip wire-illegal fields from stored Anthropic content blocks.
Older capture code emitted ``"is_error": null`` on ``tool_result`` Drops ``tool_result.is_error: null`` — the API rejects null there
blocks; the Anthropic API rejects null there (the field is optional though the field is optional.
but, when present, must be boolean). We omit the key on read so
historical rows don't break continuation.
""" """
if not isinstance(content, list): if not isinstance(content, list):
return content return content
@@ -188,27 +137,17 @@ def _sanitize_content(content: Any) -> Any:
async def rewrite_messages( async def rewrite_messages(
session: AsyncSession, *, conversation_id: int, messages: list[dict[str, Any]] session: AsyncSession, *, conversation_id: int, messages: list[dict[str, Any]]
) -> None: ) -> None:
"""Replace the conversation's stored messages with ``messages``. """Replace the conversation's stored messages (full overwrite, no branch history).
The user said no branch history — we overwrite on fork. Cheap at Deletes and flushes before inserting — SQLAlchemy's default INSERT-before-DELETE
our volume; if it ever matters we can switch to soft-delete + flush order would otherwise collide with ``uq_msg_conv_seq``.
branch pointers.
""" """
# Bulk-delete and flush before inserting the new sequence: SQLAlchemy's
# unit-of-work flushes INSERTs before DELETEs by default, which would
# trip ``uq_msg_conv_seq`` when the new rows reuse the same seq numbers
# as the soon-to-be-deleted ones.
# SQLModel descriptors resolve to ColumnElement at runtime but to bare
# ``int`` in ty's stubs; the select-path at line 135 lives behind sqlmodel's
# own ``select`` overloads that hide it, but ``sqlalchemy.delete().where``
# uses the raw stubs.
await session.execute( # ty: ignore[deprecated] await session.execute( # ty: ignore[deprecated]
delete(ConversationMessage).where( delete(ConversationMessage).where(
ConversationMessage.conversation_id == conversation_id # ty: ignore[invalid-argument-type] ConversationMessage.conversation_id == conversation_id # ty: ignore[invalid-argument-type]
) )
) )
await session.flush() await session.flush()
# Insert the new sequence.
for seq, m in enumerate(messages): for seq, m in enumerate(messages):
session.add( session.add(
ConversationMessage( ConversationMessage(
@@ -218,7 +157,6 @@ async def rewrite_messages(
content_json=json.dumps(m["content"], separators=(",", ":")), content_json=json.dumps(m["content"], separators=(",", ":")),
) )
) )
# Bump conversation.updated_at.
conv = await session.get(Conversation, conversation_id) conv = await session.get(Conversation, conversation_id)
if conv is not None: if conv is not None:
from datetime import UTC, datetime from datetime import UTC, datetime
@@ -228,21 +166,11 @@ async def rewrite_messages(
await session.commit() await session.commit()
# ---- alignment ----------------------------------------------------------
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
class _StoredDisplayTurn: class _StoredDisplayTurn:
"""A "display turn" reconstructed from stored raw messages. """A "display turn" reconstructed from stored raw messages.
``role`` is ``"user"`` (single user-prompt message) or Parser-equivalent view used to diff the file against the DB.
``"assistant"`` (one or more assistant messages, optionally
interleaved with user-only-tool_result messages). ``messages`` is
the slice of stored raw messages this display turn covers, in
order. ``spoken_text`` and ``skeleton`` are the
parser-equivalents for diff purposes; ``text_segment_count`` lets
us refuse a splice when the user edited across a tool boundary in
a way we can't safely undo.
""" """
role: str role: str
@@ -255,10 +183,8 @@ class _StoredDisplayTurn:
def _group_display_turns(stored: list[dict[str, Any]]) -> list[_StoredDisplayTurn]: def _group_display_turns(stored: list[dict[str, Any]]) -> list[_StoredDisplayTurn]:
"""Walk raw stored messages, group them into Obsidian-visible turns. """Walk raw stored messages, group them into Obsidian-visible turns.
A user-prompt message (``role=user`` with string content, or list A user-prompt message (no ``tool_result`` blocks) opens a new turn;
content with no ``tool_result`` blocks) opens a user display turn. tool-result-only messages roll into the current assistant turn.
Otherwise it's a tool-result follow-up and rolls into the current
assistant display turn.
""" """
out: list[_StoredDisplayTurn] = [] out: list[_StoredDisplayTurn] = []
i = 0 i = 0
@@ -277,7 +203,6 @@ def _group_display_turns(stored: list[dict[str, Any]]) -> list[_StoredDisplayTur
) )
i += 1 i += 1
continue continue
# Assistant display turn: collect consecutive non-prompt messages.
group: list[dict[str, Any]] = [] group: list[dict[str, Any]] = []
while i < len(stored): while i < len(stored):
m = stored[i] m = stored[i]
@@ -299,14 +224,16 @@ def _group_display_turns(stored: list[dict[str, Any]]) -> list[_StoredDisplayTur
def _is_user_prompt(content: Any) -> bool: def _is_user_prompt(content: Any) -> bool:
"""A user message is a *prompt* unless its content carries tool_result blocks.""" """A user message is a *prompt* unless its content carries tool_result blocks.
Unknown content shapes are conservatively treated as a prompt.
"""
if isinstance(content, str): if isinstance(content, str):
return True return True
if isinstance(content, list): if isinstance(content, list):
return not any( return not any(
isinstance(b, dict) and b.get("type") == "tool_result" for b in content isinstance(b, dict) and b.get("type") == "tool_result" for b in content
) )
# Unknown shape — be conservative, treat as prompt.
return True return True
@@ -328,13 +255,9 @@ def _summarize_assistant_group(
) -> tuple[str, list[str], int]: ) -> tuple[str, list[str], int]:
"""Compute (spoken_text, tool_skeleton, text_segment_count) for a display group. """Compute (spoken_text, tool_skeleton, text_segment_count) for a display group.
Mirrors what ``parser.parse_assistant_structure`` would produce when Must mirror ``parser.parse_assistant_structure``: text blocks collapse into
re-parsing the rendered version of this group: consecutive text one segment, tool_use becomes a skeleton entry, tool_result/thinking are invisible.
blocks across assistant messages collapse into one text segment;
tool_use blocks become skeleton entries; tool_result messages and
thinking blocks are invisible.
""" """
# See ``diff_and_fork`` for why the parser-type imports are deferred.
from beaver_gateway.frontends.markdown.parser import TextSegment, ToolSegment from beaver_gateway.frontends.markdown.parser import TextSegment, ToolSegment
segments: list[TextSegment | ToolSegment] = [] segments: list[TextSegment | ToolSegment] = []
@@ -351,7 +274,6 @@ def _summarize_assistant_group(
for msg in group: for msg in group:
if msg["role"] == "user": if msg["role"] == "user":
# tool_result message — boundary for text but emits no segment.
_flush() _flush()
continue continue
content = msg.get("content") content = msg.get("content")
@@ -368,7 +290,6 @@ def _summarize_assistant_group(
elif btype == "tool_use": elif btype == "tool_use":
_flush() _flush()
segments.append(ToolSegment(name=str(blk.get("name", "")))) segments.append(ToolSegment(name=str(blk.get("name", ""))))
# thinking: skip silently
_flush() _flush()
spoken_chunks = [s.text for s in segments if isinstance(s, TextSegment)] spoken_chunks = [s.text for s in segments if isinstance(s, TextSegment)]
spoken = "\n\n".join(c for c in spoken_chunks if c).strip() spoken = "\n\n".join(c for c in spoken_chunks if c).strip()
@@ -377,30 +298,14 @@ def _summarize_assistant_group(
return spoken, skeleton, text_count return spoken, skeleton, text_count
# ---- the core algorithm -------------------------------------------------
def diff_and_fork( def diff_and_fork(
*, stored: list[dict[str, Any]], incoming: list[ParsedTurn] *, stored: list[dict[str, Any]], incoming: list[ParsedTurn]
) -> ForkOutcome: ) -> ForkOutcome:
"""Align the incoming parsed file against stored history. """Align the incoming parsed file against stored history.
``stored`` is the raw Anthropic-shape message list from the DB ``incoming`` must end with a user turn (the new prompt); raises otherwise.
(one entry per ``ConversationMessage`` row). ``incoming`` is the Segment-class imports below are deferred to avoid a cycle with ``parser``.
user-visible turn list from the markdown parser. The last
``incoming`` entry must be a user turn — that's the new prompt
triggering this request.
Returns a :class:`ForkOutcome` whose ``messages`` is what the
backend should run on and whose ``persist_messages`` is the
canonical history to store in the DB once the backend's
synthesized cycle is appended.
""" """
# ``parser`` lives under ``frontends/markdown/`` whose ``__init__``
# eagerly loads ``frontend.py``, which in turn imports this module
# — pulling the parser at module-import time creates a cycle. The
# helpers below import the segment classes lazily inside their own
# function bodies to break it.
if not incoming or incoming[-1].role != "user": if not incoming or incoming[-1].role != "user":
msg = ( msg = (
"diff_and_fork expects incoming to end with a user turn " "diff_and_fork expects incoming to end with a user turn "
@@ -417,16 +322,10 @@ def diff_and_fork(
if divergence is None and len(prior_incoming) < len(stored_groups): if divergence is None and len(prior_incoming) < len(stored_groups):
if _file_lags_store(stored_groups, len(prior_incoming), new_user_turn): if _file_lags_store(stored_groups, len(prior_incoming), new_user_turn):
# Not a deletion — the file simply never received turns we
# already ran. Adopt the stored tail verbatim so history stays
# structured and its fingerprint still matches the live
# session's. See ``_file_lags_store`` for why this matters.
spliced_groups.extend( spliced_groups.extend(
list(g.messages) for g in stored_groups[len(prior_incoming) :] list(g.messages) for g in stored_groups[len(prior_incoming) :]
) )
else: else:
# Incoming truncated stored (user deleted some past turns).
# Truncate stored to match.
divergence = len(prior_incoming) divergence = len(prior_incoming)
backend_msgs, persist_msgs = _assemble_tail( backend_msgs, persist_msgs = _assemble_tail(
@@ -448,26 +347,9 @@ def _file_lags_store(
) -> bool: ) -> bool:
"""Is the shorter incoming file a stale view rather than a deletion? """Is the shorter incoming file a stale view rather than a deletion?
A file with fewer display turns than the DB has two possible causes, Tell: if the DB already holds the submitted prompt at the position the
and they need opposite handling: file stops at, it's a lagging render, not a deletion — misjudging this
forks history, breaks the fingerprint, and respawns the backend session.
* the user deleted trailing turns — we should truncate to match;
* the turn ran, was persisted, but its reply never made it back into
the ``.md`` (the render lost a race with the user's next prompt, or
the reply rendered to nothing visible). The file is simply behind.
The tell is the prompt the user is submitting right now: if the DB
already holds it at exactly the position the file stops at, this is a
re-submission of a turn we've already run, not a deletion. Nobody
deletes a turn and immediately retypes it verbatim.
Getting this wrong is expensive and self-sustaining. Forking here
flattens every post-divergence turn into plain text (losing tool_use /
tool_result structure), persists that flattened history, and changes
the conversation fingerprint — so the backend's session pool misses,
spawns a fresh ``claude``, reseeds it from a multi-MB JSONL, and
strands the previous process. The file still lags afterwards, so the
next turn does it again.
""" """
if prior_len >= len(stored_groups): if prior_len >= len(stored_groups):
return False return False
@@ -480,12 +362,8 @@ def _walk_prefix(
) -> tuple[list[list[dict[str, Any]]], int | None, bool]: ) -> tuple[list[list[dict[str, Any]]], int | None, bool]:
"""Walk incoming vs stored side-by-side until first divergence. """Walk incoming vs stored side-by-side until first divergence.
Returns the spliced/matched group list (one entry per matched An empty incoming skeleton (no tool callouts rendered) means prose
display turn, each carrying the raw messages we'll feed to the alone decides the match; a mismatch forces a fresh backend session.
backend for that turn), the divergence index (``None`` if all
of ``prior_incoming`` matched) and whether any assistant prose
was spliced in from the file - a rewritten reply keeps the
structure but must not resume the session that said otherwise.
""" """
from beaver_gateway.frontends.markdown.parser import TextSegment, ToolSegment from beaver_gateway.frontends.markdown.parser import TextSegment, ToolSegment
@@ -506,8 +384,6 @@ def _walk_prefix(
s.name for s in inc.structure if isinstance(s, ToolSegment) s.name for s in inc.structure if isinstance(s, ToolSegment)
) )
inc_text_count = sum(1 for s in inc.structure if isinstance(s, TextSegment)) inc_text_count = sum(1 for s in inc.structure if isinstance(s, TextSegment))
# Files rendered without tool callouts (§3.10) carry no skeleton:
# prose alone decides whether the turn matched.
if inc_skeleton and inc_skeleton != st.skeleton: if inc_skeleton and inc_skeleton != st.skeleton:
return spliced_groups, i, edited return spliced_groups, i, edited
if inc.text == st.spoken_text: if inc.text == st.spoken_text:
@@ -576,21 +452,10 @@ def _splice_in_place(
) -> list[dict[str, Any]] | None: ) -> list[dict[str, Any]] | None:
"""Copy the stored messages, swapping only their text block contents. """Copy the stored messages, swapping only their text block contents.
Rebuilding a turn from the file loses everything the markdown never Preserves message/thinking-block boundaries that a rebuild would lose
carried — thinking blocks, and the message boundaries claude chose. (and that a cache-hit continuation depends on); returns ``None`` if
Both matter: an ``assistant[thinking] + assistant[text]`` pair (what text blocks aren't 1:1 with incoming segments, so the caller rebuilds instead.
claude emits for a reasoning turn) collapses into a single message,
so the history is one message shorter than the one the backend
pooled its live session under, and the next turn misses the cache
and respawns. Substituting in place keeps the message count and the
invisible blocks exactly as stored.
Returns ``None`` when stored text blocks and incoming text segments
aren't one-to-one — consecutive text blocks merge into a single
rendered segment, so there'd be no way to know how to split the
edited prose back apart. The caller then rebuilds instead.
""" """
# See ``diff_and_fork`` for why the parser-type import is deferred.
from beaver_gateway.frontends.markdown.parser import TextSegment from beaver_gateway.frontends.markdown.parser import TextSegment
new_texts = [ new_texts = [
@@ -626,18 +491,9 @@ def _splice_by_rebuild(
) -> list[dict[str, Any]] | None: ) -> list[dict[str, Any]] | None:
"""Rebuild an assistant display turn with new text + stored tool_use blocks. """Rebuild an assistant display turn with new text + stored tool_use blocks.
Walks the incoming structure; for each ``TextSegment`` emits a Matches tool_use blocks to incoming ``ToolSegment``s by position; returns
text block into the current assistant message; for each ``None`` (caller forks) if a matching stored ``tool_result`` is missing.
``ToolSegment`` consumes the next stored ``tool_use`` block (by
position), closes the current assistant message, emits the
matching ``tool_result`` user message, and opens a new assistant
message. Final ``TextSegment`` closes the last assistant message.
Returns ``None`` if we can't find a matching tool_result for some
tool_use (stored history is malformed) — caller falls back to
fork.
""" """
# See ``diff_and_fork`` for why this import is deferred.
from beaver_gateway.frontends.markdown.parser import TextSegment from beaver_gateway.frontends.markdown.parser import TextSegment
tool_uses, tool_results_by_id = _harvest_tool_blocks(stored_group) tool_uses, tool_results_by_id = _harvest_tool_blocks(stored_group)
@@ -664,8 +520,6 @@ def _splice_by_rebuild(
if current_asst: if current_asst:
spliced.append({"role": "assistant", "content": current_asst}) spliced.append({"role": "assistant", "content": current_asst})
elif not spliced: elif not spliced:
# Defensive: assistant turn with no text and no tools makes no
# sense; caller will treat as fork.
return None return None
return spliced return spliced
@@ -1,13 +1,8 @@
"""Vault files for ``deep`` conversations that were not typed into a file (§3.10). """Vault files for ``deep`` conversations that were not typed into a file.
``materialize`` is the markdown frontend's answer to ``spawn(kind=deep)``: ``materialize`` creates the file for a conversation spawned elsewhere;
a new file in the vault with ``agent`` + ``conversation_id`` frontmatter ``run`` appends replies produced by other origins and keeps the canonical
and the ``(markdown, path)`` binding. ``run`` tails the gateway bus and history in sync so a continuation typed in Obsidian resumes the same session.
appends every ``reply`` of a markdown-bound conversation to its file -
the seed turn of a spawn, a message posted through ``/api``, a turn
that came in over ``/v1/messages`` - and stamps the same exchange into
the canonical history, so a continuation typed in Obsidian aligns
against the store and resumes the same SDK session instead of reseeding.
""" """
from __future__ import annotations from __future__ import annotations
+17 -100
View File
@@ -1,18 +1,7 @@
"""Parse a markdown chat file into Anthropic ``MessageParam`` history. """Parse a markdown chat file into Anthropic ``MessageParam`` history.
The file format is documented in ``frontends/markdown/__init__.py``: Turn markers are ``### User:`` / ``### Assistant:`` H3 headers; backend
``### User:`` / ``### Assistant:`` H3 headers split turns, optional input drops thinking/tool callouts, keeping assistant turns text-only.
``---`` HRs between turns are visual-only, ``> [!thinking]-`` and
``> [!tool]- <name>`` callouts mark structured assistant content.
For backend consumption we strip thinking and tool_use callouts —
assistant turns become text-only. Rationale: history replay through
claude-code's JSONL injection only needs the *narrated* answer (the
thinking signatures expire and the original tool_results aren't
captured in the renderer's output, so a faithful tool_use round-trip
isn't possible today). The renderer keeps callouts in the file because
they're informational for the human reader; the parser drops them when
shaping the backend's input.
""" """
from __future__ import annotations from __future__ import annotations
@@ -51,11 +40,7 @@ class TextSegment:
class ToolSegment: class ToolSegment:
"""A ``> [!tool]- <name>`` callout placeholder. """A ``> [!tool]- <name>`` callout placeholder.
Only the tool ``name`` is captured the " · summary" suffix on the Only the tool name is captured; the summary suffix is decorative.
callout title and the JSON body inside the quote block are
decorative for the human reader; the canonical tool_use block lives
in the DB and is keyed by *position+name* against the structure
parsed here.
""" """
name: str name: str
@@ -64,26 +49,11 @@ class ToolSegment:
AssistantSegment = TextSegment | ToolSegment AssistantSegment = TextSegment | ToolSegment
# Turn marker — must be exactly ``### User:`` or ``### Assistant:`` on
# its own line. Trailing whitespace tolerated; nothing after the colon
# on the same line (any inline content would mean the user typed
# something that just happens to look like a header, and we'd rather
# misparse than silently fold inline content into a turn).
_TURN_RE = re.compile(r"^###\s+(User|Assistant):\s*$", re.MULTILINE) _TURN_RE = re.compile(r"^###\s+(User|Assistant):\s*$", re.MULTILINE)
# Callout-start lines we strip from assistant turns when extracting
# text. We don't try to parse the contents — for backend input we just
# need to drop the whole quoted block.
_CALLOUT_START_RE = re.compile(r"^>\s+\[!(thinking|tool)\]") _CALLOUT_START_RE = re.compile(r"^>\s+\[!(thinking|tool)\]")
# Tool-callout title line: ``> [!tool]- <name>`` or ``> [!tool]- <name> · <summary>``.
# We only need the ``<name>`` part for skeleton matching; the summary is
# decorative (built by ``renderer.summarize_tool_input`` from inputs the
# user can edit visually without semantic consequence).
_TOOL_TITLE_RE = re.compile(r"^>\s+\[!tool\]-\s*(.*?)\s*$") _TOOL_TITLE_RE = re.compile(r"^>\s+\[!tool\]-\s*(.*?)\s*$")
# Renderer joins name + summary with " · " (U+00B7) — see
# ``renderer.summarize_tool_input``. We split on it to recover the
# bare tool name.
_TOOL_TITLE_SEP = " · " _TOOL_TITLE_SEP = " · "
@@ -91,14 +61,8 @@ _TOOL_TITLE_SEP = " · "
class ParsedTurn: class ParsedTurn:
"""One turn extracted from the chat file. """One turn extracted from the chat file.
``role`` is ``"user"`` or ``"assistant"``. ``text`` is the spoken ``text`` is spoken content only; ``structure`` (assistant turns only)
content with callouts stripped and HRs dropped — used both as the carries the ordered text/tool segments used to align with stored history.
backend's ``MessageParam.content`` (back-compat with the existing
parser shape) and as the diff key against stored turns.
``structure`` is non-empty only for assistant turns: an ordered
list of ``TextSegment`` / ``ToolSegment`` reflecting the visible
layout of the assistant block, used by the conversation store to
align with the canonical tool_use blocks held in DB.
""" """
role: str role: str
@@ -110,16 +74,8 @@ class ParsedTurn:
class ParsedFile: class ParsedFile:
"""Result of parsing a single chat ``.md``. """Result of parsing a single chat ``.md``.
``metadata`` is the YAML frontmatter as a plain dict (empty if the ``messages`` is text-only history for the backend; ``turns`` carries
file has none). ``messages`` is the conversation history shaped for per-turn structure; ``body`` is the raw content after frontmatter.
``Backend.complete`` — assistant turns are text-only. ``turns`` is
1:1 with ``messages`` and carries the per-turn structure (for
assistant turns) that the conversation store needs to detect
text-only edits vs. structural forks. ``body`` is the raw markdown
content *after* the frontmatter is stripped; the renderer needs it
when it appends a new assistant turn so it can preserve whatever
the human typed verbatim (including any callouts or HRs they
added).
""" """
metadata: dict[str, Any] metadata: dict[str, Any]
@@ -131,15 +87,9 @@ class ParsedFile:
def parse(text: str) -> ParsedFile: def parse(text: str) -> ParsedFile:
"""Parse a chat ``.md`` into ``(metadata, body, messages, turns)``. """Parse a chat ``.md`` into ``(metadata, body, messages, turns)``.
A file with no turn markers but non-empty body is treated as a A bare file (no turn markers) is treated as a single user turn; a
single user turn — the friendly path for "user types into a new tool-only assistant turn gets placeholder ``" "`` content since the
file and hits send" before any turn markers exist. backend rejects an empty string.
Assistant turns that have *only* tool callouts (no spoken text) are
preserved here even though their ``MessageParam.content`` is empty
— the structure carries tool-segment information the conversation
store needs for skeleton matching. The renderer in practice always
emits at least a trailing text block, so this branch is defensive.
""" """
parsed = frontmatter.loads(text) parsed = frontmatter.loads(text)
metadata = dict(parsed.metadata) metadata = dict(parsed.metadata)
@@ -175,14 +125,6 @@ def parse(text: str) -> ParsedFile:
) )
) )
elif has_tools: elif has_tools:
# Tool-only assistant turn: nothing to feed the backend
# as ``content`` (it'd reject an empty string), but the
# structure must survive so the store can align it
# against stored tool_use blocks. We synthesize a
# single-space text content for backend round-trip; the
# conversation store will replace this payload with the
# canonical stored blocks before the backend ever sees
# it on a continuation.
messages.append({"role": "assistant", "content": " "}) messages.append({"role": "assistant", "content": " "})
parsed_turns.append( parsed_turns.append(
ParsedTurn(role="assistant", text="", structure=tuple(structure)) ParsedTurn(role="assistant", text="", structure=tuple(structure))
@@ -196,18 +138,8 @@ def parse(text: str) -> ParsedFile:
def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]: def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]:
"""Walk an assistant turn body, return its ordered text/tool segments. """Walk an assistant turn body, return its ordered text/tool segments.
Tool callouts become :class:`ToolSegment` with just the tool name — Tool callouts become :class:`ToolSegment` (name only); thinking callouts
the title's optional ``" · summary"`` suffix and the JSON body and HR lines are stripped; boundary-empty text segments are dropped.
inside the quote block are decorative; the canonical tool_use
block is held in the conversation store. Thinking callouts are
stripped entirely (they were never round-trippable through the
file — signatures expire). HR separator lines drop out.
Empty / whitespace-only text segments at the boundaries (start,
end, between adjacent tool callouts) are dropped so the skeleton
is robust against renderer whitespace choices; a non-empty text
segment with surrounding whitespace is trimmed on both ends but
preserved.
""" """
segments: list[TextSegment | ToolSegment] = [] segments: list[TextSegment | ToolSegment] = []
pending_text: list[str] = [] pending_text: list[str] = []
@@ -216,9 +148,6 @@ def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]:
if not pending_text: if not pending_text:
return return
joined = "\n".join(pending_text) joined = "\n".join(pending_text)
# Collapse runs of >2 blank lines (created when we stripped a
# mid-block callout) into one so the diff against a re-render
# is stable.
cleaned = re.sub(r"\n{3,}", "\n\n", joined).strip() cleaned = re.sub(r"\n{3,}", "\n\n", joined).strip()
pending_text.clear() pending_text.clear()
if cleaned: if cleaned:
@@ -231,7 +160,6 @@ def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]:
callout_match = _CALLOUT_START_RE.match(line) callout_match = _CALLOUT_START_RE.match(line)
if callout_match: if callout_match:
kind = callout_match.group(1) kind = callout_match.group(1)
# Capture tool name *before* advancing past the block.
if kind == "tool": if kind == "tool":
title_match = _TOOL_TITLE_RE.match(line) title_match = _TOOL_TITLE_RE.match(line)
title = title_match.group(1) if title_match else "" title = title_match.group(1) if title_match else ""
@@ -239,9 +167,7 @@ def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]:
_flush_text() _flush_text()
segments.append(ToolSegment(name=name)) segments.append(ToolSegment(name=name))
else: else:
# Thinking callout — drop the whole block, emit nothing.
_flush_text() _flush_text()
# Skip the rest of the quote block.
while i < len(lines) and lines[i].lstrip().startswith(">"): while i < len(lines) and lines[i].lstrip().startswith(">"):
i += 1 i += 1
continue continue
@@ -257,11 +183,7 @@ def parse_assistant_structure(raw: str) -> list[TextSegment | ToolSegment]:
def _segments_to_spoken_text(segments: list[TextSegment | ToolSegment]) -> str: def _segments_to_spoken_text(segments: list[TextSegment | ToolSegment]) -> str:
r"""Reduce a structure list to the spoken-text view the backend sees. r"""Reduce a structure list to the spoken-text view the backend sees.
Concatenates :class:`TextSegment` contents with ``\n\n`` between Concatenates text segments with ``\n\n``, dropping tool segments.
them, dropping :class:`ToolSegment` entries. Equivalent to what
the pre-Conversation-store parser did — we keep that behavior so
existing fingerprints (frontmatter ``fingerprint`` field) stay
valid.
""" """
chunks = [s.text for s in segments if isinstance(s, TextSegment)] chunks = [s.text for s in segments if isinstance(s, TextSegment)]
return "\n\n".join(c for c in chunks if c).strip() return "\n\n".join(c for c in chunks if c).strip()
@@ -290,15 +212,11 @@ def resolve_agent(
return default return default
# ---- internals ---------------------------------------------------------
def _split_turns(body: str) -> list[tuple[str, str]]: def _split_turns(body: str) -> list[tuple[str, str]]:
"""Walk turn markers, return ``[(role_lc, raw_body), ...]``. """Walk turn markers, return ``[(role_lc, raw_body), ...]``.
Body for each turn is everything between this marker and the next A marker is an exact ``### User:`` / ``### Assistant:`` line; trailing
(or EOF). Leading marker line itself is dropped. We don't trim content after the colon means it's not a marker, not a turn boundary.
whitespace here — that's per-role.
""" """
matches = list(_TURN_RE.finditer(body)) matches = list(_TURN_RE.finditer(body))
if not matches: if not matches:
@@ -315,9 +233,8 @@ def _split_turns(body: str) -> list[tuple[str, str]]:
def _strip_hrs(raw: str) -> str: def _strip_hrs(raw: str) -> str:
"""Drop decorative ``---`` separator lines (whole-line HRs only). """Drop decorative ``---`` separator lines (whole-line HRs only).
A ``---`` mid-paragraph (rare, but possible) stays. Only lines that Only lines that are exactly ``---`` (with optional surrounding
are *exactly* the HR after optional surrounding whitespace are whitespace) are removed; a ``---`` mid-paragraph stays.
removed — those are the ones the renderer emits between turns.
""" """
lines = raw.splitlines() lines = raw.splitlines()
kept = [ln for ln in lines if ln.strip() != "---"] kept = [ln for ln in lines if ln.strip() != "---"]
@@ -1,10 +1,7 @@
"""Render Anthropic ``Message`` (and individual user turns) into markdown. """Render Anthropic ``Message`` (and user turns) into markdown.
The renderer is one-way: it produces the human-facing artifact in the One-way: produces the human-facing artifact only. The parser strips
vault. The parser strips tool/thinking callouts when reshaping the file tool/thinking callouts separately when reshaping history for backend replay.
for backend replay — so what we write here is purely for the human
reader (and for the cross-frontend logger, which materializes turns
from other frontends).
""" """
from __future__ import annotations from __future__ import annotations
@@ -30,23 +27,11 @@ __all__ = [
] ]
# Empty ``### User:`` block appended after each assistant reply so the
# human has an obvious place to type the next turn. Parser drops empty
# user blocks, so this doesn't re-trigger dispatch on its own.
# Blank line after the header, like every rendered turn - the file stays
# symmetric whether the human or the gateway wrote the marker.
USER_SCAFFOLD = "### User:\n\n" USER_SCAFFOLD = "### User:\n\n"
# Default 4-backtick fence so tool results that contain literal ```` ``` ````
# don't collide. JSON inputs use 3 backticks because they almost never
# contain ``` and we get language syntax highlighting in Obsidian for free.
FENCE = "````" FENCE = "````"
# Input keys we dangle after the tool name in the callout title, best
# first — purely cosmetic. ``description`` wins because when a tool
# offers one it's a human-written summary of the call, which beats a
# truncated shell command or path.
_TITLE_KEYS = ("description", "path", "file", "filename", "url", "command", "query") _TITLE_KEYS = ("description", "path", "file", "filename", "url", "command", "query")
@@ -61,24 +46,15 @@ def render_assistant_text(text: str) -> str:
def render_assistant_message(message: Message) -> str: def render_assistant_message(message: Message) -> str:
"""Render an assistant ``Message`` (with content blocks) into a turn block. """Render an assistant ``Message`` into a turn block.
Blocks render in their original order: Tool-use blocks render to nothing; only text and thinking content reach
the file (see :mod:`.parser` for how the reverse strip works).
* ``ThinkingBlock`` → ``> [!thinking]-`` collapsed callout
* ``TextBlock`` → plain text (the spoken answer)
* ``ToolUseBlock`` → nothing (§3.10: tool calls never reach the file;
"what the agent is doing" is the activity panel fed by SSE)
Blank lines separate adjacent blocks; trailing newline guarantees
the next ``---`` / ``### User:`` marker lands on its own line.
""" """
parts: list[str] = ["### Assistant:", ""] parts: list[str] = ["### Assistant:", ""]
for block in message.content: for block in message.content:
lines = list(_render_block(block)) lines = list(_render_block(block))
if not lines: if not lines:
# Tool calls render to nothing - no separator for them either,
# or every tool leaves a blank line behind.
continue continue
parts.extend(lines) parts.extend(lines)
parts.append("") parts.append("")
@@ -88,19 +64,12 @@ def render_assistant_message(message: Message) -> str:
def render_user_param(param: MessageParam) -> str: def render_user_param(param: MessageParam) -> str:
"""Render a ``MessageParam`` user message into a ``### User:`` block. """Render a ``MessageParam`` user message into a ``### User:`` block.
Used by the cross-frontend logger when materializing turns from Tool_result blocks in the content list are dropped silently.
other frontends. Tool_result blocks in the content list are dropped
silently — the markdown view doesn't track them (see ``parser.py``).
""" """
content = param.get("content", "") content = param.get("content", "")
if isinstance(content, str): if isinstance(content, str):
text = content text = content
else: else:
# The Anthropic SDK types ``content`` as a union of typed-dict
# *Param classes plus pydantic block models — both shapes appear
# in practice (raw incoming JSON yields dicts, SDK-built params
# yield BaseModels). Treat each entry as a dict-like and pull
# ``text`` opportunistically.
chunks = [ chunks = [
str(blk.get("text", "")) str(blk.get("text", ""))
for blk in content for blk in content
@@ -111,11 +80,9 @@ def render_user_param(param: MessageParam) -> str:
def append_to_body(existing: str, new_block: str) -> str: def append_to_body(existing: str, new_block: str) -> str:
"""Append ``new_block`` to ``existing`` with a decorative HR separator. """Append ``new_block`` to ``existing`` with a decorative ``---`` separator.
Preserves the original body verbatim (whitespace, callouts, any The separator is visual only; the parser ignores it.
formatting the human added). The HR is purely visual: parser ignores
it.
""" """
head = existing.rstrip() head = existing.rstrip()
if head: if head:
@@ -124,18 +91,9 @@ def append_to_body(existing: str, new_block: str) -> str:
def summarize_tool_input(name: str, tool_input: object) -> str: def summarize_tool_input(name: str, tool_input: object) -> str:
"""Build the ``[!tool]- <summary>`` title string. """Build the ``[!tool]- <summary>`` title, picking one salient input field."""
Tries to pick a single salient field (``description``, ``path``,
``command``, etc.) from the input dict so the collapsed callout shows
something meaningful in Obsidian. Falls back to just the tool name.
"""
if not isinstance(tool_input, dict): if not isinstance(tool_input, dict):
return name return name
# Anthropic ``ToolUseBlock.input`` is typed as ``object`` — the
# SDK's runtime value is always a JSON dict (str→Any), so a local
# cast keeps the rest of the function readable without sprinkling
# per-line type narrowing on every ``.get`` call.
d = cast("dict[str, Any]", tool_input) d = cast("dict[str, Any]", tool_input)
for key in _TITLE_KEYS: for key in _TITLE_KEYS:
value = d.get(key) value = d.get(key)
@@ -145,9 +103,6 @@ def summarize_tool_input(name: str, tool_input: object) -> str:
return name return name
# ---- internals ---------------------------------------------------------
def _render_block(block: object) -> Iterable[str]: def _render_block(block: object) -> Iterable[str]:
if isinstance(block, TextBlock): if isinstance(block, TextBlock):
text = (block.text or "").strip() text = (block.text or "").strip()
@@ -157,7 +112,6 @@ def _render_block(block: object) -> Iterable[str]:
if isinstance(block, ThinkingBlock): if isinstance(block, ThinkingBlock):
yield from _render_thinking(block.thinking or "") yield from _render_thinking(block.thinking or "")
return return
# Tool-use blocks and unknown block types never reach the file.
def _render_thinking(text: str) -> Iterable[str]: def _render_thinking(text: str) -> Iterable[str]:
@@ -167,12 +121,7 @@ def _render_thinking(text: str) -> Iterable[str]:
def adaptive_fence(content: str) -> str: def adaptive_fence(content: str) -> str:
"""Return a backtick fence at least one longer than the longest run in ``content``. """Return a backtick fence longer than any backtick run in ``content``."""
Currently unused (tool *results* aren't persisted yet) — kept here
so when result capture lands the rendering side already has the
primitive.
"""
longest = 0 longest = 0
for match in re.finditer(r"`+", content): for match in re.finditer(r"`+", content):
longest = max(longest, len(match.group(0))) longest = max(longest, len(match.group(0)))
+31 -83
View File
@@ -1,36 +1,7 @@
"""External MCP frontend (Phase 3.1). """External MCP frontend.
A streamable-HTTP gateway in front of the internal MCP aggregator Streamable-HTTP gateway that reverse-proxies the internal MCP aggregator,
(``beaver_gateway.mcp.internal_app``). The aggregator hosts every adding bearer auth, an audit log line per request, and a discovery page.
declared ``McpServer`` (``python_tool``, stdio proxy, HTTP proxy)
under ``/mcp/<name>`` plus a flat ``/mcp/all`` bundle on
``127.0.0.1:INTERNAL_MCP_PORT`` — that's the *internal* shape.
This frontend re-exposes those namespaces on its own port directly at
``/<name>/`` (no ``/mcp/`` prefix in the external routes — the port
itself already disambiguates). Caddy / nginx / Cloudflare in front
typically strips a prefix back on: ``domain.com/mcp/* → :8001/*``,
controlled by the operator's reverse-proxy config and surfaced to the
admin dashboard via ``public_base_url``. Three additions on top of the
raw aggregator:
* **Bearer auth** — ``Authorization: Bearer <token>``, ``X-Api-Key``,
or ``?token=<…>`` query string. All three forms verify against the
same :class:`TokenStore` as ``AnthropicMessagesFrontend``.
* **Audit log** — one line per request (token name, namespace,
request method/path, response status). The DB-backed audit log lives
in Phase 4; for now we just emit a structured log line.
* **Discovery page** at ``GET /`` (auth-gated) — HTML rendered with a
tiny inline Jinja2 template listing every namespace plus copy-pastable
config snippets for Cursor / claude.ai / Claude Desktop.
Why a reverse proxy and not a second mount? FastMCP's session managers
are tied to the lifespan they were created in; running the same
aggregator under two uvicorn servers double-initializes state. Building
two parallel aggregators would double upstream connections (two
subprocesses for every stdio MCP, two HTTP clients for every remote).
A loopback proxy keeps one source of truth — the internal aggregator —
and lets us layer policy on the outside.
""" """
from __future__ import annotations from __future__ import annotations
@@ -61,10 +32,6 @@ if TYPE_CHECKING:
_log = logging.getLogger("beaver_gateway.frontends.mcp_server") _log = logging.getLogger("beaver_gateway.frontends.mcp_server")
# Hop-by-hop headers that must NOT be forwarded across an HTTP proxy
# (RFC 7230 §6.1). Bypassing this filter would break chunked transfer
# encoding when ``Content-Length`` arrives, or upstream-aware proxies
# would refuse the second hop's connection-pool reuse.
_HOP_BY_HOP_HEADERS = frozenset( _HOP_BY_HOP_HEADERS = frozenset(
{ {
"connection", "connection",
@@ -80,11 +47,6 @@ _HOP_BY_HOP_HEADERS = frozenset(
} }
) )
# Standard auth-bearing headers we *do not* forward to the internal app —
# the internal app is on loopback with no auth of its own, and forwarding
# the inbound bearer would only confuse it. Each method-specific MCP
# request from the upstream Cursor/etc. carries a fresh ``mcp-session-id``
# that we *must* forward.
_AUTH_HEADERS = frozenset({"authorization", "x-api-key"}) _AUTH_HEADERS = frozenset({"authorization", "x-api-key"})
@@ -119,17 +81,16 @@ class McpServerFrontend(Frontend):
self._http = None self._http = None
def _build_app(self, runtime: GatewayRuntime) -> Starlette: # noqa: ARG002 def _build_app(self, runtime: GatewayRuntime) -> Starlette: # noqa: ARG002
"""Build the Starlette app.
Literal routes (``/``, ``/healthz``) are listed before the
namespace wildcard so they win the match instead of being
swallowed by it; two routes per namespace cover both ``/x`` and
``/x/y`` since Starlette won't fold them into one.
"""
routes = [ routes = [
Route("/", self._discovery, methods=["GET"]), Route("/", self._discovery, methods=["GET"]),
Route("/healthz", self._healthz, methods=["GET"]), Route("/healthz", self._healthz, methods=["GET"]),
# Namespaces mount at the root of this port — the port
# itself already disambiguates this from any other gateway
# surface. Two routes per namespace so both the
# trailing-slash and sub-path forms work (``/time`` AND
# ``/time/foo``); Starlette doesn't fold them into one
# route automatically. The literal routes above (``/``,
# ``/healthz``) are listed first and win the match, so
# they're not eaten by ``/{namespace}``.
Route( Route(
"/{namespace}", "/{namespace}",
self._proxy_endpoint, self._proxy_endpoint,
@@ -151,7 +112,7 @@ class McpServerFrontend(Frontend):
token_name, err = await _verify_request(request, runtime) token_name, err = await _verify_request(request, runtime)
if err is not None: if err is not None:
return err return err
assert token_name is not None # noqa: S101 — narrow for ty assert token_name is not None # noqa: S101
base = external_base(request, runtime) base = external_base(request, runtime)
html = _render_discovery_page( html = _render_discovery_page(
base_url=base, namespaces=list(runtime.mcps), actor=token_name base_url=base, namespaces=list(runtime.mcps), actor=token_name
@@ -165,7 +126,7 @@ class McpServerFrontend(Frontend):
token_name, err = await _verify_request(request, runtime) token_name, err = await _verify_request(request, runtime)
if err is not None: if err is not None:
return err return err
assert token_name is not None # noqa: S101 — narrow for ty assert token_name is not None # noqa: S101
namespace = request.path_params["namespace"] namespace = request.path_params["namespace"]
subpath = request.path_params.get("path", "") subpath = request.path_params.get("path", "")
@@ -201,11 +162,12 @@ class McpServerFrontend(Frontend):
) )
def _upstream_url(self, namespace: str, subpath: str) -> str | None: def _upstream_url(self, namespace: str, subpath: str) -> str | None:
"""Resolve ``namespace`` to its internal loopback URL.
``ALL_NAMESPACE`` isn't in the URL map, so its URL is synthesized
from any per-domain entry's authority.
"""
runtime = self._require_runtime() runtime = self._require_runtime()
# ``all`` is built by the aggregator unconditionally when at least
# one MCP is configured; the URL map only contains per-domain
# entries (see ``build_internal_app``), so we synthesize ``all``'s
# loopback URL from any per-domain URL's authority.
if namespace == ALL_NAMESPACE: if namespace == ALL_NAMESPACE:
sample = next(iter(runtime.mcp_internal_urls.values()), None) sample = next(iter(runtime.mcp_internal_urls.values()), None)
if sample is None: if sample is None:
@@ -232,13 +194,8 @@ async def _verify_request(
) -> tuple[str | None, JSONResponse | None]: ) -> tuple[str | None, JSONResponse | None]:
"""Accept ``Authorization: Bearer``, ``X-Api-Key``, or ``?token=``. """Accept ``Authorization: Bearer``, ``X-Api-Key``, or ``?token=``.
The third form is the escape hatch for clients that can only put Returns ``(actor_name, None)`` on success, else ``(None, error_response)``
secrets in the URL (claude.ai's MCP config historically did this). — 401 for a missing/unknown token, 403 for one whose scope doesn't cover this call.
All three roads end at the same :class:`TokenStore`. Returns
``(actor_name, None)`` on success, ``(None, 401|403)`` otherwise
— the caller forwards the response as-is. Splitting auth vs scope
failures matters: 401 says "send me a token", 403 says "this token
is real but not for this endpoint".
""" """
api_key = request.headers.get("x-api-key") api_key = request.headers.get("x-api-key")
if api_key: if api_key:
@@ -276,9 +233,7 @@ def _forbidden(scope: str, required: str) -> JSONResponse:
def _join_subpath(base_url: str, subpath: str) -> str: def _join_subpath(base_url: str, subpath: str) -> str:
"""Concatenate the loopback URL with the proxied sub-path. """Concatenate the loopback URL with the proxied sub-path.
``base_url`` always ends in ``/`` (the aggregator publishes URLs ``base_url`` always ends in ``/`` to avoid Starlette's redirect dance.
that way to avoid Starlette's 307 redirect dance); the sub-path is
appended verbatim, with the query string handled by the caller.
""" """
if subpath: if subpath:
return base_url + subpath.lstrip("/") return base_url + subpath.lstrip("/")
@@ -294,18 +249,14 @@ async def _reverse_proxy(
actor: str, actor: str,
runtime: GatewayRuntime, runtime: GatewayRuntime,
) -> StreamingResponse | JSONResponse: ) -> StreamingResponse | JSONResponse:
"""Bidirectionally stream an MCP request between client internal aggregator. """Bidirectionally stream an MCP request between client and internal aggregator.
Streamable-HTTP MCP responses can be a long-running SSE stream The audit row is written right after the upstream response headers
(tools that emit partial progress) or a one-shot JSON body; we arrive, before relaying its body, so a client-truncated stream is
don't peek — just relay chunks as they arrive in either direction still audited.
until both sides close.
""" """
qs = request.url.query qs = request.url.query
if qs: if qs:
# Drop ``?token=`` from the forwarded URL — internal app doesn't
# need it, and propagating creds further than necessary widens
# the leak surface (logs, metrics, traces all see query strings).
scrubbed = _scrub_query(qs, drop={"token"}) scrubbed = _scrub_query(qs, drop={"token"})
if scrubbed: if scrubbed:
upstream_url = f"{upstream_url}?{scrubbed}" upstream_url = f"{upstream_url}?{scrubbed}"
@@ -351,9 +302,6 @@ async def _reverse_proxy(
request.url.path, request.url.path,
upstream_resp.status, upstream_resp.status,
) )
# Audit at upstream-response time: status reflects the MCP call's
# outcome (200 / tool-error / 4xx). Streaming relay below may be
# cut short by the client, but the row is already in by then.
await audit.log( await audit.log(
runtime, runtime,
actor=f"token:{actor}", actor=f"token:{actor}",
@@ -370,7 +318,6 @@ async def _reverse_proxy(
async for chunk in upstream_resp.content.iter_any(): async for chunk in upstream_resp.content.iter_any():
yield chunk yield chunk
except (aiohttp.ClientError, asyncio.CancelledError): except (aiohttp.ClientError, asyncio.CancelledError):
# Caller hung up or upstream dropped — just stop relaying.
return return
finally: finally:
upstream_resp.release() upstream_resp.release()
@@ -390,6 +337,11 @@ async def _request_body_iter(request: Request) -> AsyncIterator[bytes]:
def _forward_headers(request: Request) -> dict[str, str]: def _forward_headers(request: Request) -> dict[str, str]:
"""Drop hop-by-hop headers (RFC 7230 §6.1) and inbound auth headers.
The internal aggregator is loopback-only with no auth of its own,
so forwarding the caller's bearer would only confuse it.
"""
out: dict[str, str] = {} out: dict[str, str] = {}
for key, value in request.headers.items(): for key, value in request.headers.items():
lowered = key.lower() lowered = key.lower()
@@ -422,11 +374,7 @@ def _scrub_query(query: str, *, drop: frozenset[str] | set[str]) -> str:
def _render_discovery_page(*, base_url: str, namespaces: list[Any], actor: str) -> str: def _render_discovery_page(*, base_url: str, namespaces: list[Any], actor: str) -> str:
"""Render the auth-gated namespace + config-snippet page. """Render the auth-gated namespace + config-snippet page."""
Inline HTML (no Jinja file) — keeps Phase 3 free of template-dir
plumbing that Phase 4's AdminFrontend will own.
"""
name_list = [getattr(ns, "name", str(ns)) for ns in namespaces] name_list = [getattr(ns, "name", str(ns)) for ns in namespaces]
rows = ( rows = (
"\n".join( "\n".join(
@@ -460,7 +408,7 @@ def _escape(value: str) -> str:
) )
_DISCOVERY_TEMPLATE = "\n".join( # noqa: FLY002 — readability beats one-string-blob _DISCOVERY_TEMPLATE = "\n".join( # noqa: FLY002
[ [
"<!doctype html>", "<!doctype html>",
'<html lang="en">', '<html lang="en">',
+1 -1
View File
@@ -46,7 +46,7 @@ def build_root_app(
] ]
for fe in mounted: for fe in mounted:
app = fe.app() app = fe.app()
assert app is not None # noqa: S101 - filtered above; narrows for ty assert app is not None # noqa: S101
assert fe.path is not None # noqa: S101 assert fe.path is not None # noqa: S101
routes.append(Mount(fe.path, app=app, name=fe.name or fe.path.strip("/"))) routes.append(Mount(fe.path, app=app, name=fe.name or fe.path.strip("/")))
for path, app in (extra or {}).items(): for path, app in (extra or {}).items():
+2 -4
View File
@@ -29,10 +29,8 @@ async def events_with_heartbeat(
) -> AsyncIterator[Any]: ) -> AsyncIterator[Any]:
"""Pass ``events`` through, yielding ``None`` after ``interval`` seconds of silence. """Pass ``events`` through, yielding ``None`` after ``interval`` seconds of silence.
One in-flight ``__anext__`` task is reused across timeouts: a second Only one consumer may iterate the result at a time; cancelling the
consumer on the same async generator raises ``RuntimeError``. outer scope cancels the in-flight upstream fetch instead of leaking it.
Cancellation of the outer scope cancels that task instead of leaving
it dangling.
""" """
src = events.__aiter__() src = events.__aiter__()
next_task: asyncio.Task[Any] | None = None next_task: asyncio.Task[Any] | None = None
@@ -1,13 +1,8 @@
"""One ``sendMessageDraft`` stream per running turn (§3.8). """One ``sendMessageDraft`` stream per running turn.
The client folds a draft into the message that follows only when their Telegram folds a draft into the following message only when their texts are
texts are identical - so the last push is the final text itself, rendered identical, so the last push is the final text verbatim. Drafts live 30s and
exactly as the outbox will send it, with no status line. throttle to about one edit/s, so this refreshes on a timer, not every delta.
A draft is ephemeral and lives 30 s, Telegram throttles edits to about one
per second per chat, and thinking or a tool call would otherwise look like a
hang - so the draft opens with a status line straight away, is refreshed on
a timer rather than on every delta, and is kept alive while nothing changes.
""" """
from __future__ import annotations from __future__ import annotations
@@ -1,13 +1,8 @@
"""``TelegramFrontend`` - the private chat with the bot as the window (§3.8). """``TelegramFrontend`` - the private chat with the bot as the window.
A private chat with topics has no General: the gateway makes one topic for A private chat with topics has no General, so the gateway makes and rebinds
the master (``master_topic``) and rebinds it to every new master; any other one topic for the master; any other topic is a branch, and a message into a
topic is a branch. The user makes a topic and the new one spawns it. Replies stream as drafts and land through the outbox.
first message in it spawns the branch (``seed=morning``); a message into a
topic whose branch is merged or closed spawns a new branch on the same
topic. Replies stream as drafts and land through the outbox; turns that
came from other windows are mirrored with a marker; ``origin=system`` is
never shown. ``AskUserQuestion`` becomes inline buttons (§3.7).
""" """
from __future__ import annotations from __future__ import annotations
@@ -62,13 +57,10 @@ _COMMANDS = ("merge", "new", "chat", "status", "help", "start")
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
class Attachments: class Attachments:
"""Where files from Telegram go. """Where files land.
Files land in ``<root>/YYYY-MM-DD/<unixts>-<name>`` (the day in ``tz``) ``ephemeral`` uses the gateway's data dir, ``vault`` an inbox under
and whatever is older than ``keep_days`` is swept; ``None`` never ``dir``; ``keep_days`` sweeps older files, ``None`` never sweeps.
sweeps. ``ephemeral`` - under the gateway's data dir; ``vault`` -
``dir`` is an inbox inside the agent's zone: the agent moves keepers
next to the note, the rest is swept after ``keep_days``.
""" """
mode: Literal["ephemeral", "vault"] = "ephemeral" mode: Literal["ephemeral", "vault"] = "ephemeral"
@@ -148,8 +140,6 @@ class TelegramFrontend(Frontend):
self._reactions: dict[int, tuple[int, int]] = {} self._reactions: dict[int, tuple[int, int]] = {}
self._tasks: set[asyncio.Task[None]] = set() self._tasks: set[asyncio.Task[None]] = set()
# ---- Frontend --------------------------------------------------------
def agent_for(self, kind: Kind) -> str | None: def agent_for(self, kind: Kind) -> str | None:
return {"master": self.master_agent, "branch": self.branch_agent}.get(kind) return {"master": self.master_agent, "branch": self.branch_agent}.get(kind)
@@ -223,8 +213,6 @@ class TelegramFrontend(Frontend):
self._topic_names[target[1]] = f"{prefix}{name}" self._topic_names[target[1]] = f"{prefix}{name}"
return True return True
# ---- plumbing --------------------------------------------------------
@property @property
def bot(self) -> Bot: def bot(self) -> Bot:
if self._bot is None: if self._bot is None:
@@ -355,8 +343,6 @@ class TelegramFrontend(Frontend):
binding=(FRONTEND, self._ext(thread_id)), binding=(FRONTEND, self._ext(thread_id)),
) )
# ---- inbox -----------------------------------------------------------
async def _handle(self, update: Update) -> None: async def _handle(self, update: Update) -> None:
if update.message is not None: if update.message is not None:
await self._on_message(update.message) await self._on_message(update.message)
@@ -556,8 +542,6 @@ class TelegramFrontend(Frontend):
if folder.is_dir() and not any(folder.iterdir()): if folder.is_dir() and not any(folder.iterdir()):
folder.rmdir() folder.rmdir()
# ---- commands --------------------------------------------------------
async def _command( async def _command(
self, command: str, args: str, message: Message, thread_id: int | None self, command: str, args: str, message: Message, thread_id: int | None
) -> None: ) -> None:
@@ -663,8 +647,6 @@ class TelegramFrontend(Frontend):
self._tasks.add(task) self._tasks.add(task)
task.add_done_callback(self._tasks.discard) task.add_done_callback(self._tasks.discard)
# ---- bus -------------------------------------------------------------
async def _events(self) -> None: async def _events(self) -> None:
async for event in self.bus.stream(): async for event in self.bus.stream():
try: try:
@@ -790,8 +772,6 @@ class TelegramFrontend(Frontend):
) )
await self._deliver(conv, text, turn_id=turn_id, key=f"{turn_id}:reply") await self._deliver(conv, text, turn_id=turn_id, key=f"{turn_id}:reply")
# ---- drafts ------------------------------------------------------------
async def _open_draft( async def _open_draft(
self, key: str, event: Event, target: tuple[int, int | None] self, key: str, event: Event, target: tuple[int, int | None]
) -> None: ) -> None:
@@ -841,8 +821,6 @@ class TelegramFrontend(Frontend):
if draft is not None: if draft is not None:
await draft.stop() await draft.stop()
# ---- questions (§3.7) ---------------------------------------------------
async def _ask( async def _ask(
self, conv: Conversation, event: Event, target: tuple[int, int | None] self, conv: Conversation, event: Event, target: tuple[int, int | None]
) -> None: ) -> None:
@@ -1,8 +1,7 @@
"""Long-polling inbox (§3.8). """Long-polling inbox.
Every update lands in ``telegram_updates`` before the offset moves past it; Every update lands in ``telegram_updates`` before the offset moves past it;
a worker handles rows from the table, oldest first, and finishes whatever a a worker handles rows from the table, oldest first, resuming after a restart.
previous process left unprocessed at startup.
""" """
from __future__ import annotations from __future__ import annotations
@@ -1,8 +1,7 @@
"""Outbox (§3.8): a reply is a ``deliveries`` row first, a message second. """Outbox: a reply is a ``deliveries`` row first, a message second.
Rows are sent oldest first, retried with backoff on network errors and Rows are sent oldest first, retried with backoff on network errors, resent
flood limits, resent as plain text when Telegram rejects our HTML, and as plain text when Telegram rejects our HTML, and given up once it's gone.
given up only when Telegram says the window is gone.
""" """
from __future__ import annotations from __future__ import annotations
+8 -46
View File
@@ -1,16 +1,7 @@
"""Cross-frontend turn record. """Cross-frontend turn record.
Frontends that finish a turn (the Anthropic Messages frontend, the Frontends emit a :class:`TurnRecord` to every handler in
markdown frontend) emit a :class:`TurnRecord` to every handler in ``GatewayRuntime.turn_log_handlers`` after a turn finishes.
``GatewayRuntime.turn_log_handlers``. The markdown frontend uses this
to persist chats from other frontends into the Obsidian vault see
``frontends/markdown/crossfront.py``.
Kept tiny on purpose: it carries the structured-enough payload a logger
needs (which agent ran, input history, the assembled assistant reply)
and nothing else. The full event stream is gone by the time handlers
run if a future consumer needs deltas it would subscribe at a lower
level, not here.
""" """
from __future__ import annotations from __future__ import annotations
@@ -26,10 +17,6 @@ if TYPE_CHECKING:
__all__ = ["TurnRecord", "slugify"] __all__ = ["TurnRecord", "slugify"]
# Filesystem-safe slug: collapse anything that isn't a word char or
# space/hyphen to a hyphen, then squash runs of separators. Aimed at
# letting users build filenames from ``record.first_user_text`` without
# hand-rolling sanitization in every config.
_SLUG_BAD_RE = re.compile(r"[^\w\s\-]+", flags=re.UNICODE) _SLUG_BAD_RE = re.compile(r"[^\w\s\-]+", flags=re.UNICODE)
_SLUG_SEP_RE = re.compile(r"[\s\-]+", flags=re.UNICODE) _SLUG_SEP_RE = re.compile(r"[\s\-]+", flags=re.UNICODE)
@@ -37,10 +24,7 @@ _SLUG_SEP_RE = re.compile(r"[\s\-]+", flags=re.UNICODE)
def slugify(text: str, *, maxlen: int = 40) -> str: def slugify(text: str, *, maxlen: int = 40) -> str:
"""Sanitize ``text`` for use as a filename fragment. """Sanitize ``text`` for use as a filename fragment.
Strips punctuation, collapses whitespace/hyphens into single ``-``, Truncates to ``maxlen`` and returns ``"untitled"`` for empty input.
and truncates to ``maxlen``. Returns ``"untitled"`` for empty input.
Unicode letters are preserved (Obsidian handles them fine; macOS
and modern Linux fs's too).
""" """
cleaned = _SLUG_BAD_RE.sub(" ", text).strip() cleaned = _SLUG_BAD_RE.sub(" ", text).strip()
cleaned = _SLUG_SEP_RE.sub("-", cleaned).strip("-") cleaned = _SLUG_SEP_RE.sub("-", cleaned).strip("-")
@@ -55,17 +39,8 @@ def slugify(text: str, *, maxlen: int = 40) -> str:
class TurnRecord: class TurnRecord:
"""One completed turn, as seen by a frontend. """One completed turn, as seen by a frontend.
``input_messages`` is the conversation history sent to the backend ``input_messages`` excludes the assistant reply; ``source`` names the
(everything *before* the assistant reply). ``output_message`` is the producing frontend so a cross-frontend logger can skip its own turns.
finalized assistant ``Message`` (post-accumulation, with all content
blocks attached). ``system`` is the per-request system prompt if any
agents own their own ``system_prompt``, this is the override the
caller passed; most handlers can ignore it.
``source`` names the frontend that produced the record so the
cross-frontend logger can avoid logging its own turns (markdown
frontend writing the file would otherwise also receive its own
broadcast and double-write).
""" """
agent_name: str agent_name: str
@@ -76,12 +51,7 @@ class TurnRecord:
@property @property
def first_user_text(self) -> str: def first_user_text(self) -> str:
"""Plain text of the *earliest* user turn in this conversation. """Empty string if the input history has no user turn."""
Useful for naming new files by topic. Empty string if the input
history somehow has no user turn (shouldn't happen — turns are
broadcast only after a user assistant cycle).
"""
for msg in self.input_messages: for msg in self.input_messages:
if msg.get("role") == "user": if msg.get("role") == "user":
return _text_of(msg.get("content", "")) return _text_of(msg.get("content", ""))
@@ -89,7 +59,6 @@ class TurnRecord:
@property @property
def last_user_text(self) -> str: def last_user_text(self) -> str:
"""Plain text of the most recent user turn (the trigger)."""
for msg in reversed(self.input_messages): for msg in reversed(self.input_messages):
if msg.get("role") == "user": if msg.get("role") == "user":
return _text_of(msg.get("content", "")) return _text_of(msg.get("content", ""))
@@ -109,10 +78,8 @@ class TurnRecord:
def _text_of(content: object) -> str: def _text_of(content: object) -> str:
"""Flatten an Anthropic ``MessageParam.content`` field to a plain string. """Flatten an Anthropic ``MessageParam.content`` field to a plain string.
Handles both shapes the SDK accepts: raw string, or a list of block Handles both raw-string and block-list shapes; unrecognized content is
dicts (we pull ``text`` blocks only). Anything we don't recognize is silently skipped since this exists for naming, not faithful reconstruction.
silently skipped these helpers exist for naming files, not for
faithful content reconstruction.
""" """
if isinstance(content, str): if isinstance(content, str):
return content return content
@@ -121,11 +88,6 @@ def _text_of(content: object) -> str:
for blk in content: for blk in content:
if not isinstance(blk, dict): if not isinstance(blk, dict):
continue continue
# ``MessageParam.content`` is typed as a union of typed-dicts
# per Anthropic SDK; we only care about plain ``text`` blocks
# and look them up via ``Any`` to dodge the keyed-typed-dict
# variance gymnastics (``ty`` won't let an open dict alias a
# closed-keyed one).
d: Any = blk d: Any = blk
if d.get("type") == "text": if d.get("type") == "text":
parts.append(str(d.get("text", ""))) parts.append(str(d.get("text", "")))
+1
View File
@@ -0,0 +1 @@
"""Cron, webhook and event jobs, deferred injects, the subscription budget."""
+126
View File
@@ -0,0 +1,126 @@
"""What a job is: its triggers, the budget it respects, and what a run may do."""
from __future__ import annotations
import logging
from dataclasses import dataclass
from datetime import datetime, timedelta
from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from collections.abc import Awaitable, Callable, Coroutine
from beaver_gateway.conversations.distill import LineCap
from beaver_gateway.conversations.injects import Priority
from beaver_gateway.conversations.service import Conversations, DistillResult
from beaver_gateway.jobs.scheduler import Scheduler
from beaver_gateway.storage.models import Conversation
__all__ = ["Budget", "Job", "JobRun"]
_log = logging.getLogger(__name__)
@dataclass(frozen=True, slots=True)
class Job:
name: str
run: Callable[[JobRun], Awaitable[None]]
cron: str | None = None
webhook: bool = False
events: tuple[str, ...] = ()
critical: bool = True
dedupe: bool = True
@property
def entrypoint(self) -> str:
return f"job:{self.name}"
@dataclass(frozen=True, slots=True)
class Budget:
threshold: float = 0.7
tokens: int | None = None
window: timedelta = timedelta(hours=5)
limit_window: str = "five_hour"
@dataclass(slots=True)
class JobRun:
job: Job
trigger: str
payload: dict[str, Any]
scheduler: Scheduler
@property
def conversations(self) -> Conversations:
return self.scheduler.conversations
async def master(self) -> Conversation | None:
masters = await self.conversations.find(kind="master", status="open", limit=1)
return masters[0] if masters else None
async def inject_master(
self, text: str, *, urgency: Priority = "normal", origin: str | None = None
) -> bool:
master = await self.master()
if master is None:
_log.error(
"job %s: no open master, inject lost: %s", self.job.name, text[:200]
)
return False
await self.conversations.inject(
master, text, urgency=urgency, origin=origin or self.job.name
)
return True
async def spawn_job(
self,
*,
agent: str,
text: str,
title: str | None = None,
line_cap: LineCap | None = None,
) -> Conversation:
"""A headless job turn; ``line_cap`` bounces a rewrite past the cap."""
return await self.conversations.spawn(
kind="job",
agent=agent,
seed="brief",
text=text,
title=title,
origin="job",
flags={"line_cap": line_cap.as_flags()} if line_cap else None,
)
async def close_idle(
self,
*,
kind: str = "deep",
days: int = 2,
limit: int = 3,
since: datetime | None = None,
) -> list[DistillResult]:
"""Close chats quiet for ``days``, at most ``limit`` per run."""
out: list[DistillResult] = []
for conv in await self.conversations.idle(
kind=kind, days=days, since=since, limit=limit
):
try:
out.append(
await self.conversations.distill(conv, reason=f"idle {days}d")
)
except Exception: # noqa: BLE001
_log.exception("closing idle %s failed", conv.external_id)
return out
async def retry_in(self, delay: timedelta) -> None:
await self.scheduler.trigger(
self.job, self.payload, delay=delay, trigger=self.trigger
)
async def rotate(self) -> list[Conversation]:
rotation = self.scheduler.rotation
return await rotation.tick() if rotation is not None else []
def background(self, coro: Coroutine[Any, Any, Any]) -> None:
self.scheduler.background(coro)
+5 -130
View File
@@ -1,13 +1,7 @@
"""Jobs and deferred injects on pgqueuer (§3.6, §4.5). """Jobs and deferred injects on pgqueuer.
A job is a name, a handler and its triggers: a cron expression, the A job is a name, a handler and its triggers (cron, webhook, bus events); a
webhook ``/hooks/<name>``, gateway bus events. The executor is pgqueuer on handler only queues work, and the turn itself runs in the conversation's worker.
the gateway's own Postgres (a dedicated autocommit connection for
LISTEN/NOTIFY), so cron ticks, webhook deliveries and the ``schedule``
tool's one-off injects all live in one table and survive a restart. A
handler only queues work for a conversation and returns; the turn itself
runs in the conversation's worker. Non-critical jobs step aside while the
subscription window is past its threshold.
""" """
from __future__ import annotations from __future__ import annotations
@@ -34,6 +28,7 @@ from starlette.responses import JSONResponse
from starlette.routing import Route from starlette.routing import Route
from beaver_gateway.conversations.service import parse_at from beaver_gateway.conversations.service import parse_at
from beaver_gateway.jobs.job import Budget, Job, JobRun
from beaver_gateway.storage.models import JobRunRecord from beaver_gateway.storage.models import JobRunRecord
if TYPE_CHECKING: if TYPE_CHECKING:
@@ -44,10 +39,9 @@ if TYPE_CHECKING:
from pgqueuer.ports.driver import Driver from pgqueuer.ports.driver import Driver
from starlette.requests import Request from starlette.requests import Request
from beaver_gateway.conversations.distill import LineCap
from beaver_gateway.conversations.injects import Priority from beaver_gateway.conversations.injects import Priority
from beaver_gateway.conversations.rotation import Rotation from beaver_gateway.conversations.rotation import Rotation
from beaver_gateway.conversations.service import Conversations, DistillResult from beaver_gateway.conversations.service import Conversations
from beaver_gateway.storage.models import Conversation from beaver_gateway.storage.models import Conversation
__all__ = ["INJECT", "Budget", "Job", "JobRun", "LocalCron", "Scheduler", "next_run"] __all__ = ["INJECT", "Budget", "Job", "JobRun", "LocalCron", "Scheduler", "next_run"]
@@ -73,111 +67,6 @@ class LocalCron(ScheduleExecutor):
return next_run(self.parameters.expression, self.tz) return next_run(self.parameters.expression, self.tz)
@dataclass(frozen=True, slots=True)
class Job:
name: str
run: Callable[[JobRun], Awaitable[None]]
cron: str | None = None
webhook: bool = False
events: tuple[str, ...] = ()
critical: bool = True
dedupe: bool = True
@property
def entrypoint(self) -> str:
return f"job:{self.name}"
@dataclass(frozen=True, slots=True)
class Budget:
threshold: float = 0.7
tokens: int | None = None
window: timedelta = timedelta(hours=5)
limit_window: str = "five_hour"
@dataclass(slots=True)
class JobRun:
job: Job
trigger: str
payload: dict[str, Any]
scheduler: Scheduler
@property
def conversations(self) -> Conversations:
return self.scheduler.conversations
async def master(self) -> Conversation | None:
masters = await self.conversations.find(kind="master", status="open", limit=1)
return masters[0] if masters else None
async def inject_master(
self, text: str, *, urgency: Priority = "normal", origin: str | None = None
) -> bool:
master = await self.master()
if master is None:
_log.error(
"job %s: no open master, inject lost: %s", self.job.name, text[:200]
)
return False
await self.conversations.inject(
master, text, urgency=urgency, origin=origin or self.job.name
)
return True
async def spawn_job(
self,
*,
agent: str,
text: str,
title: str | None = None,
line_cap: LineCap | None = None,
) -> Conversation:
"""A headless job turn; ``line_cap`` bounces a rewrite past the cap."""
return await self.conversations.spawn(
kind="job",
agent=agent,
seed="brief",
text=text,
title=title,
origin="job",
flags={"line_cap": line_cap.as_flags()} if line_cap else None,
)
async def close_idle(
self,
*,
kind: str = "deep",
days: int = 2,
limit: int = 3,
since: datetime | None = None,
) -> list[DistillResult]:
"""§4.5: close chats quiet for ``days``, at most ``limit`` per run."""
out: list[DistillResult] = []
for conv in await self.conversations.idle(
kind=kind, days=days, since=since, limit=limit
):
try:
out.append(
await self.conversations.distill(conv, reason=f"idle {days}d")
)
except Exception: # noqa: BLE001
_log.exception("closing idle %s failed", conv.external_id)
return out
async def retry_in(self, delay: timedelta) -> None:
await self.scheduler.trigger(
self.job, self.payload, delay=delay, trigger=self.trigger
)
async def rotate(self) -> list[Conversation]:
rotation = self.scheduler.rotation
return await rotation.tick() if rotation is not None else []
def background(self, coro: Coroutine[Any, Any, Any]) -> None:
self.scheduler.background(coro)
class Scheduler: class Scheduler:
def __init__( def __init__(
self, self,
@@ -212,8 +101,6 @@ class Scheduler:
def job(self, name: str) -> Job | None: def job(self, name: str) -> Job | None:
return self._jobs.get(name) return self._jobs.get(name)
# ---- lifecycle -----------------------------------------------------
async def start(self) -> None: async def start(self) -> None:
if self._driver is not None: if self._driver is not None:
queries = Queries(self._driver) queries = Queries(self._driver)
@@ -274,8 +161,6 @@ class Scheduler:
self._dispatch(job, trigger="event", payload=dict(event)) self._dispatch(job, trigger="event", payload=dict(event))
) )
# ---- dispatch ------------------------------------------------------
async def _dispatch( async def _dispatch(
self, job: Job, *, trigger: str, payload: dict[str, Any] self, job: Job, *, trigger: str, payload: dict[str, Any]
) -> None: ) -> None:
@@ -350,8 +235,6 @@ class Scheduler:
) )
return int(ids[0]) if ids and ids[0] is not None else None return int(ids[0]) if ids and ids[0] is not None else None
# ---- deferred injects ----------------------------------------------
async def schedule( async def schedule(
self, self,
conv: Conversation, conv: Conversation,
@@ -424,8 +307,6 @@ class Scheduler:
origin="schedule", origin="schedule",
) )
# ---- budget --------------------------------------------------------
async def utilization(self) -> float | None: async def utilization(self) -> float | None:
now = datetime.now(UTC) now = datetime.now(UTC)
values: list[float] = [] values: list[float] = []
@@ -446,8 +327,6 @@ class Scheduler:
utilization = await self.utilization() utilization = await self.utilization()
return utilization is not None and utilization > self.budget.threshold return utilization is not None and utilization > self.budget.threshold
# ---- introspection -------------------------------------------------
async def snapshot(self) -> dict[str, Any]: async def snapshot(self) -> dict[str, Any]:
crons: dict[str, PgSchedule] = {} crons: dict[str, PgSchedule] = {}
queue: list[dict[str, Any]] = [] queue: list[dict[str, Any]] = []
@@ -534,8 +413,6 @@ class Scheduler:
"payload": row.payload, "payload": row.payload,
} }
# ---- http ----------------------------------------------------------
def app(self, authorize: Callable[[Request], Awaitable[Any]]) -> Starlette: def app(self, authorize: Callable[[Request], Awaitable[Any]]) -> Starlette:
async def hook(request: Request) -> JSONResponse: async def hook(request: Request) -> JSONResponse:
await authorize(request) await authorize(request)
@@ -549,8 +426,6 @@ class Scheduler:
return Starlette(routes=[Route("/{name}", hook, methods=["POST"])]) return Starlette(routes=[Route("/{name}", hook, methods=["POST"])])
# ---- internals -----------------------------------------------------
def background(self, coro: Coroutine[Any, Any, Any]) -> None: def background(self, coro: Coroutine[Any, Any, Any]) -> None:
self._spawn(coro) self._spawn(coro)
+1 -1
View File
@@ -1,4 +1,4 @@
"""MCP server definitions and (later) the internal aggregator app.""" """MCP server definitions and the internal aggregator app."""
from __future__ import annotations from __future__ import annotations
+2 -6
View File
@@ -1,11 +1,7 @@
"""Proxy ``FastMCP`` servers for user-declared external MCPs (``stdio``/``http``). """Proxy ``FastMCP`` servers for user-declared external MCPs (``stdio``/``http``).
Both flavours end up as a ``FastMCPProxy`` instance, built via Built via ``fastmcp.server.create_proxy``; the proxy lazily opens its
``fastmcp.server.create_proxy``. The proxy lazily opens the underlying transport on first request, so unused MCPs cost nothing until routed to.
client transport when the first MCP request arrives, so we don't pay
for connections that nothing routes to. From the aggregator app's
point of view a proxy is indistinguishable from a regular ``FastMCP``
namespace same ``http_app`` surface, same mount semantics.
""" """
from __future__ import annotations from __future__ import annotations
+19 -43
View File
@@ -1,26 +1,7 @@
"""Internal MCP aggregator — one ASGI app, N FastMCP namespaces + ``all``. """Internal MCP aggregator — one ASGI app, N FastMCP namespaces + ``all``.
Each ``McpServer`` declared in the user's config becomes its own Each ``McpServer`` in the user's config is mounted under ``/mcp/<name>``;
``FastMCP`` instance (regular for ``python_tool``, ``FastMCPProxy`` for ``/mcp/all/`` unions every namespace's tools for single-MCP clients.
``stdio``/``http``) and is mounted under ``/mcp/<name>`` on a single
Starlette app. This app runs on ``127.0.0.1:INTERNAL_MCP_PORT`` (not
EXPOSE'd in Docker) so the ClaudeCode subprocess can reach each
namespace via loopback as a distinct MCP server URL preserving
per-domain framing while costing only one process worth of RAM
(PRD §6).
Phase 3 adds ``/mcp/all/``: a single FastMCP whose tools are the union
of every namespace's tools, prefixed by FastMCP's ``namespace_<tool>``
convention (e.g. ``time_current_time``). It's the escape-hatch for
clients that can only configure one MCP server discouraged for tool-
heavy setups (PRD §6 cites the ~95%~71% tool-selection drop on flat
namespaces) but real and reachable.
The aggregator returns both the app and a ``{name: url}`` map; Phase
``ClaudeSdkBackend`` plugs the map directly into
``BackendOptions.mcp_servers``. ``/mcp/all/`` is NOT included in that
map claude-code-agents always get per-domain framing; only the
external MCP frontend (Phase 3.1) reverse-proxies the flat endpoint.
""" """
from __future__ import annotations from __future__ import annotations
@@ -52,11 +33,13 @@ def build_internal_app(
) -> tuple[Starlette, dict[str, str], dict[str, FastMCP]]: ) -> tuple[Starlette, dict[str, str], dict[str, FastMCP]]:
"""Build the aggregator ``Starlette`` app, per-namespace URL map, and server map. """Build the aggregator ``Starlette`` app, per-namespace URL map, and server map.
``host``/``port`` only flavour the URL strings handed back actually ``host``/``port`` only flavour the URL strings handed back; the
listening on them is the caller's job (``cli.main`` runs a uvicorn caller (``cli.main``) is what actually listens on them. Published
server in a TaskGroup). We accept the address here so callers don't URLs carry a trailing slash so ``Mount`` doesn't 307-redirect the
have to format the URLs themselves and risk drifting from the first request. Each
``/mcp/<name>`` convention. namespace gets its own :class:`RedactingMiddleware` so a direct
``call_tool`` on a child (e.g. the Raycast backend does that) is
filtered too, not just requests through the mounted app.
Returns: Returns:
* Starlette app to serve via uvicorn. * Starlette app to serve via uvicorn.
@@ -70,10 +53,6 @@ def build_internal_app(
MCP tools into the Raycast wire). MCP tools into the Raycast wire).
""" """
servers: dict[str, FastMCP] = {spec.name: _build_server(spec) for spec in mcps} servers: dict[str, FastMCP] = {spec.name: _build_server(spec) for spec in mcps}
# The one place every tool result passes through on its way to the
# model. Attached per namespace rather than once at the top so that
# a direct ``call_tool`` on a child (the Raycast backend does that)
# is filtered too.
for server in servers.values(): for server in servers.values():
server.add_middleware(RedactingMiddleware()) server.add_middleware(RedactingMiddleware())
@@ -82,8 +61,6 @@ def build_internal_app(
} }
routes = [Mount(f"/mcp/{name}", app=app) for name, app in child_apps.items()] routes = [Mount(f"/mcp/{name}", app=app) for name, app in child_apps.items()]
# /mcp/all — flat-namespace bundle. Skip when there's nothing to
# bundle so we don't pay for an empty session manager lifecycle.
all_app = None all_app = None
if servers: if servers:
all_server = _build_all_server(servers) all_server = _build_all_server(servers)
@@ -93,11 +70,11 @@ def build_internal_app(
@asynccontextmanager @asynccontextmanager
async def lifespan(_parent: Starlette) -> AsyncIterator[None]: async def lifespan(_parent: Starlette) -> AsyncIterator[None]:
# Each FastMCP http_app stores its session manager init in its """Enter every child app's lifespan.
# own lifespan. Without entering them the streamable-HTTP layer
# 500s on every request. AsyncExitStack composes them so all Without it the streamable-HTTP layer 500s on every request;
# children come up together and unwind in reverse order on children unwind in reverse order.
# shutdown. """
async with AsyncExitStack() as stack: async with AsyncExitStack() as stack:
for child in child_apps.values(): for child in child_apps.values():
await stack.enter_async_context(child.router.lifespan_context(child)) await stack.enter_async_context(child.router.lifespan_context(child))
@@ -108,23 +85,22 @@ def build_internal_app(
yield yield
app = Starlette(routes=routes, lifespan=lifespan) app = Starlette(routes=routes, lifespan=lifespan)
# Trailing slash on the published URL skips Starlette's
# 307 redirect from ``/mcp/<name>`` to ``/mcp/<name>/`` that
# ``Mount`` produces when a child route lives at ``/``.
urls = {name: f"http://{host}:{port}/mcp/{name}/" for name in servers} urls = {name: f"http://{host}:{port}/mcp/{name}/" for name in servers}
return app, urls, servers return app, urls, servers
def _build_server(spec: McpServerT) -> FastMCP: def _build_server(spec: McpServerT) -> FastMCP:
"""Dispatch on the discriminated union to the matching builder.""" """Dispatch on the discriminated union to the matching builder.
The final branch is unreachable while ``McpServerT`` stays closed;
it exists to keep type-narrowing honest if a variant is ever added.
"""
if isinstance(spec, PythonToolMcp): if isinstance(spec, PythonToolMcp):
return build_python_tool_server(spec) return build_python_tool_server(spec)
if isinstance(spec, StdioMcp): if isinstance(spec, StdioMcp):
return build_stdio_proxy(spec) return build_stdio_proxy(spec)
if isinstance(spec, HttpMcp): if isinstance(spec, HttpMcp):
return build_http_proxy(spec) return build_http_proxy(spec)
# `McpServerT` is a closed union; this is unreachable but keeps
# type-narrowing honest if a new variant lands without updates here.
msg = f"unsupported McpServer variant: {type(spec).__name__}" msg = f"unsupported McpServer variant: {type(spec).__name__}"
raise TypeError(msg) raise TypeError(msg)
+12 -25
View File
@@ -1,22 +1,8 @@
"""Tolerant transports for upstream MCPs that don't strictly speak JSON-RPC. """Tolerant transports for upstream MCPs that don't strictly speak JSON-RPC.
Some real-world MCP servers print non-JSON chatter to stdout before / between Some MCP servers print non-JSON chatter to stdout (banners, dependency-load
their actual JSON-RPC frames (``Processing...``, banners, dependency-load messages) that the reference client forwards as exceptions, which bleed into
messages, etc.). The reference ``mcp.client.stdio.stdio_client`` parses every client UIs as warnings. This transport drops those lines instead.
stdout line as JSON-RPC and ships any parse failure downstream as an
exception, which the MCP ``ClientSession`` then logs as a warning that bleeds
into client UIs (Cursor, Cline) when they connect through us.
``LenientStdioTransport`` re-implements the stdio-client wiring with one
behavioural change: lines that don't parse as JSON-RPC are *silently
dropped* (one ``DEBUG`` log entry, no exception forwarded). Downstream
consumers see only valid messages. We keep the rest of the contract identical
to the reference client, including the spec-mandated graceful shutdown
sequence (close stdin wait SIGTERM SIGKILL).
The transport plugs into ``fastmcp.server.create_proxy`` just like
``StdioTransport`` does, so the rest of the aggregator doesn't need to
know which flavour it got.
""" """
from __future__ import annotations from __future__ import annotations
@@ -70,13 +56,15 @@ class LenientStdioTransport(ClientTransport):
cwd: str | None = None, cwd: str | None = None,
log_file: TextIO | None = None, log_file: TextIO | None = None,
) -> None: ) -> None:
"""``log_file`` takes an already-open ``TextIO``.
Unlike the upstream transport, this one does not open a ``Path``
for you.
"""
self.command = command self.command = command
self.args = args self.args = args
self.env = env self.env = env
self.cwd = cwd self.cwd = cwd
# TextIO only — pre-open Path callers themselves. The upstream
# ``StdioTransport`` opens Path for you, but we keep this thin so
# the type contract stays narrow and easy to validate.
self.log_file = log_file self.log_file = log_file
@contextlib.asynccontextmanager @contextlib.asynccontextmanager
@@ -100,7 +88,7 @@ class LenientStdioTransport(ClientTransport):
@contextlib.asynccontextmanager @contextlib.asynccontextmanager
async def _lenient_stdio_client( # noqa: PLR0915 — mirrors mcp.client.stdio.stdio_client async def _lenient_stdio_client( # noqa: PLR0915
server: StdioServerParameters, errlog: TextIO = sys.stderr server: StdioServerParameters, errlog: TextIO = sys.stderr
) -> AsyncIterator[ ) -> AsyncIterator[
tuple[ tuple[
@@ -112,8 +100,8 @@ async def _lenient_stdio_client( # noqa: PLR0915 — mirrors mcp.client.stdio.s
All differences from upstream live in ``stdout_reader``: lines that fail All differences from upstream live in ``stdout_reader``: lines that fail
``JSONRPCMessage.model_validate_json`` are logged at DEBUG and skipped, ``JSONRPCMessage.model_validate_json`` are logged at DEBUG and skipped,
never forwarded as exceptions. This is what makes warning-noisy MCPs never forwarded as exceptions. Shutdown still follows the MCP spec
quiet from the consumer's point of view. sequence: close stdin, wait, SIGTERM, then SIGKILL.
""" """
read_stream_writer, read_stream = anyio.create_memory_object_stream[ read_stream_writer, read_stream = anyio.create_memory_object_stream[
SessionMessage | Exception SessionMessage | Exception
@@ -159,7 +147,7 @@ async def _lenient_stdio_client( # noqa: PLR0915 — mirrors mcp.client.stdio.s
continue continue
try: try:
message = types.JSONRPCMessage.model_validate_json(stripped) message = types.JSONRPCMessage.model_validate_json(stripped)
except Exception: # noqa: BLE001 — by design, see module doc except Exception: # noqa: BLE001
_log.debug( _log.debug(
"lenient stdio: dropped non-JSON line: %r", "lenient stdio: dropped non-JSON line: %r",
stripped[:200], stripped[:200],
@@ -192,7 +180,6 @@ async def _lenient_stdio_client( # noqa: PLR0915 — mirrors mcp.client.stdio.s
try: try:
yield read_stream, write_stream yield read_stream, write_stream
finally: finally:
# MCP spec stdio shutdown: close stdin → wait → SIGTERM → SIGKILL.
if process.stdin: if process.stdin:
with contextlib.suppress(Exception): with contextlib.suppress(Exception):
await process.stdin.aclose() await process.stdin.aclose()
+9 -26
View File
@@ -1,26 +1,7 @@
"""Redact credentials on the way out of an MCP tool. """Redact credentials in MCP tool output before it reaches the model.
A tool result is a wider channel than the log. It goes straight into the Covers every MCP route; in-process tools like ``Bash``/``Read`` are
model's context, from there into the turn record in Postgres, and from guarded separately by ``agents.policy``.
there into the markdown transcript in ``💬 чаты`` which Obsidian Sync
carries off the machine. One ``komodo`` deploy returns the resolved
compose file, ``environment:`` block and all, so a single call can put
every secret of a stack into all four places at once.
One filter, not a list of exceptions: every MCP the model can reach is
built into a ``FastMCP`` by :mod:`beaver_gateway.mcp.internal_app`
``python_tool`` bundles like komodo, stdio subprocesses, remote HTTP
servers and every route into one of them runs its middleware chain.
That covers the per-namespace ``/mcp/<name>/`` mounts claude-code talks
to, the ``/mcp/all`` bundle, the external MCP frontend reverse-proxying
into both, and the Raycast backend's direct ``call_tool``. A new MCP in
``config.py`` is covered the day it is added, without anyone
remembering to list it here.
What this does not reach: tools that never touch a FastMCP server the
gateway's own ``gateway`` tools, and everything claude-code runs inside
its own process (``Bash``, ``Read``). Those are guarded by
:mod:`beaver_gateway.agents.policy` and the vault mounts instead.
""" """
from __future__ import annotations from __future__ import annotations
@@ -69,6 +50,12 @@ class RedactingMiddleware(Middleware):
context: MiddlewareContext[mt.CallToolRequestParams], context: MiddlewareContext[mt.CallToolRequestParams],
call_next: CallNext[mt.CallToolRequestParams, ToolResult], call_next: CallNext[mt.CallToolRequestParams, ToolResult],
) -> ToolResult: ) -> ToolResult:
"""Mask the result.
A non-``None`` ``meta`` on the masked copy takes the
``CallToolResult`` path that skips output-schema validation,
since masked structured content may no longer match it.
"""
result = await call_next(context) result = await call_next(context)
content = [_redact_block(block) for block in result.content] content = [_redact_block(block) for block in result.content]
structured = redact_data(result.structured_content) structured = redact_data(result.structured_content)
@@ -77,9 +64,5 @@ class RedactingMiddleware(Middleware):
return ToolResult( return ToolResult(
content=content, content=content,
structured_content=structured, structured_content=structured,
# Masked structured content no longer has to satisfy the
# tool's output schema; a non-None meta takes the
# ``CallToolResult`` path that skips that validation, the
# same trick ``ResponseLimitingMiddleware`` uses.
meta=result.meta if result.meta is not None else {}, meta=result.meta if result.meta is not None else {},
) )
+18 -18
View File
@@ -6,7 +6,7 @@ discriminated-union members so downstream code can ``match`` on ``kind``.
from __future__ import annotations from __future__ import annotations
from collections.abc import Callable # noqa: TC003 — runtime use by pydantic from collections.abc import Callable # noqa: TC003
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, Annotated, Literal from typing import TYPE_CHECKING, Annotated, Literal
@@ -20,39 +20,35 @@ class _BaseMcp(BaseModel):
model_config = ConfigDict(frozen=True, arbitrary_types_allowed=True) model_config = ConfigDict(frozen=True, arbitrary_types_allowed=True)
name: str name: str
"""The namespace this MCP server is mounted under."""
class StdioMcp(_BaseMcp): class StdioMcp(_BaseMcp):
"""Subprocess MCP server we spawn and connect to over stdio. """Subprocess MCP server spawned and connected to over stdio."""
``lenient`` switches the upstream stdio reader to a tolerant variant
that silently drops non-JSON-RPC lines from the subprocess's stdout
(``Processing...``-style chatter, banners, dependency-load messages).
The reference ``mcp.client.stdio.stdio_client`` forwards those parse
failures as exceptions, which bleed into Cursor/Cline UIs as warnings
when they connect through us. Default ``False`` keeps the strict
contract flip it on per-namespace for known-noisy upstreams.
"""
kind: Literal["stdio"] = "stdio" kind: Literal["stdio"] = "stdio"
command: tuple[str, ...] command: tuple[str, ...]
"""Argv used to spawn the subprocess."""
env: dict[str, str] | None = None env: dict[str, str] | None = None
"""Extra environment variables for the subprocess."""
cwd: Path | None = None cwd: Path | None = None
"""Working directory to spawn the subprocess in."""
lenient: bool = False lenient: bool = False
"""Tolerate non-JSON-RPC stdout lines instead of raising (the strict
reader's parse errors surface as warnings in Cursor/Cline). Off by default."""
class HttpMcp(_BaseMcp): class HttpMcp(_BaseMcp):
"""Remote MCP server reached over streamable HTTP. """Remote MCP server reached over streamable HTTP."""
``headers`` are forwarded on every request to the upstream MCP handy
for upstreams that authenticate via custom header rather than a Bearer
token (``auth``).
"""
kind: Literal["http"] = "http" kind: Literal["http"] = "http"
url: str url: str
"""Streamable-HTTP endpoint URL."""
auth: str | None = None auth: str | None = None
"""Bearer token sent as the upstream's ``Authorization`` header."""
headers: dict[str, str] | None = None headers: dict[str, str] | None = None
"""Extra headers forwarded on every request; for upstreams that
authenticate via a custom header instead of ``auth``."""
class PythonToolMcp(_BaseMcp): class PythonToolMcp(_BaseMcp):
@@ -60,13 +56,14 @@ class PythonToolMcp(_BaseMcp):
kind: Literal["python_tool"] = "python_tool" kind: Literal["python_tool"] = "python_tool"
tools: tuple[Callable[..., object], ...] tools: tuple[Callable[..., object], ...]
"""The callables to expose as MCP tools."""
McpServerT = Annotated[StdioMcp | HttpMcp | PythonToolMcp, Field(discriminator="kind")] McpServerT = Annotated[StdioMcp | HttpMcp | PythonToolMcp, Field(discriminator="kind")]
class McpServer: class McpServer:
"""Factory facade matching the PRD-documented config surface.""" """Factory facade for declaring MCP servers in config."""
@classmethod @classmethod
def stdio( def stdio(
@@ -78,6 +75,7 @@ class McpServer:
cwd: Path | str | None = None, cwd: Path | str | None = None,
lenient: bool = False, lenient: bool = False,
) -> StdioMcp: ) -> StdioMcp:
"""Declare a subprocess MCP server spawned over stdio."""
return StdioMcp( return StdioMcp(
name=name, name=name,
command=tuple(command), command=tuple(command),
@@ -95,10 +93,12 @@ class McpServer:
auth: str | None = None, auth: str | None = None,
headers: dict[str, str] | None = None, headers: dict[str, str] | None = None,
) -> HttpMcp: ) -> HttpMcp:
"""Declare a remote MCP server reached over streamable HTTP."""
return HttpMcp(name=name, url=url, auth=auth, headers=headers) return HttpMcp(name=name, url=url, auth=auth, headers=headers)
@classmethod @classmethod
def python_tool( def python_tool(
cls, *, name: str, tools: Iterable[Callable[..., object]] cls, *, name: str, tools: Iterable[Callable[..., object]]
) -> PythonToolMcp: ) -> PythonToolMcp:
"""Declare a namespace of Python callables exposed as MCP tools."""
return PythonToolMcp(name=name, tools=tuple(tools)) return PythonToolMcp(name=name, tools=tuple(tools))
+2 -3
View File
@@ -1,8 +1,7 @@
"""Wrap a ``PythonToolMcp`` spec into a mountable ``FastMCP`` instance. """Wrap a ``PythonToolMcp`` spec into a mountable ``FastMCP`` instance.
Each ``python_tool`` McpServer in the user's config becomes a separate Each ``python_tool`` McpServer becomes its own ``FastMCP`` namespace so
``FastMCP`` namespace one domain, one server URL so models keep the models keep the per-domain tool framing they were trained on.
per-domain framing they were trained on (see PRD §6).
""" """
from __future__ import annotations from __future__ import annotations
+1
View File
@@ -0,0 +1 @@
"""Bearer tokens, the audit log and credential redaction."""
+11 -39
View File
@@ -1,32 +1,7 @@
"""Single entry point for writing :class:`AuditLog` rows. """Single entry point for writing :class:`AuditLog` rows.
Every frontend ends up needing the same three-line pattern open a DB ``log()`` is fire-and-forget: it awaits the DB write so callers get
session, append a row, swallow failures so the user-visible request ordering, but never raises failures are logged and swallowed.
still succeeds. Phase 4.3 inlined that pattern in the admin frontend
under a private ``_audit()`` helper; Phase 4.4 lifts it here so the
Messages and MCP frontends can call the same function and so the
swallow-and-log policy lives in one place.
The contract:
* ``log(runtime, actor=..., kind=...)`` is fire-and-forget. It awaits
the DB write (so callers can ``await`` it before responding and get
ordering), but never raises if the audit insert fails, the function
emits an ``exception`` log line and returns.
* ``actor`` is a free-form string. By convention: ``"token:<name>"``
for bearer-authenticated traffic, ``"admin:<user>"`` for admin-UI
actions, ``"anon"`` for failed-auth paths we still want to record.
* ``kind`` is a short tag see :data:`KNOWN_KINDS` for the set the
current frontends emit; new tags don't need a code change here, the
column is free-form.
* ``**detail`` is JSON-serialised by :func:`append_audit`. Keep it
small: paths, methods, status codes not request bodies. Anything
passed here lands in ``AuditLog.detail_json`` verbatim.
Why a thin wrapper rather than ``append_audit`` directly: callers want
"write if you can, otherwise carry on", and pulling the try/except into
every frontend was already starting to drift (admin had it, bearer
frontends would have copy-pasted). One module, one policy.
""" """
from __future__ import annotations from __future__ import annotations
@@ -43,15 +18,11 @@ if TYPE_CHECKING:
_log = logging.getLogger("beaver_gateway.audit") _log = logging.getLogger("beaver_gateway.audit")
# Tags currently emitted by the gateway. The set is informational —
# ``AuditLog.kind`` is free-form so new code can introduce new tags
# without touching this list — but listing them here gives the admin UI
# and any downstream log consumers one canonical reference.
KNOWN_KINDS: frozenset[str] = frozenset( KNOWN_KINDS: frozenset[str] = frozenset(
{ {
"messages", # POST /v1/messages accepted "messages",
"mcp_call", # /mcp/<ns>/... proxied "mcp_call",
"tool_call", # a model's tool call seen by the PreToolUse hook "tool_call",
"login_ok", "login_ok",
"login_failed", "login_failed",
"logout", "logout",
@@ -59,6 +30,7 @@ KNOWN_KINDS: frozenset[str] = frozenset(
"token_revoke", "token_revoke",
} }
) )
"""Kinds the gateway currently emits; ``AuditLog.kind`` stays free-form."""
async def log( async def log(
@@ -69,12 +41,12 @@ async def log(
agent_name: str | None = None, agent_name: str | None = None,
**detail: Any, **detail: Any,
) -> None: ) -> None:
"""Best-effort audit insert. Never raises. """Best-effort audit insert. Never raises; DB failures are logged.
Opens its own short-lived :class:`AsyncSession` so callers don't ``actor`` is free-form (``"token:<name>"``, ``"admin:<user>"``,
have to thread one through. If the DB hiccups (table missing, ``"anon"``); ``kind`` is a short tag (see :data:`KNOWN_KINDS`); and
disk full, connection drop), we log and move on the audit trail ``**detail`` is JSON-serialised into ``AuditLog.detail_json`` keep
is observability, not a hard precondition for serving the request. it small, not full request bodies.
""" """
try: try:
async with runtime.db.session() as session: async with runtime.db.session() as session:
+36 -83
View File
@@ -1,32 +1,7 @@
"""Bearer-token verification (Phase 4.2 — DB-backed with in-memory cache). """Bearer-token verification: a DB-backed store with an in-memory cache.
The store is fed by two sources: Optional ``BOOTSTRAP_TOKENS`` env entries are checked first. Verify takes
a raw token or ``Authorization`` header and returns a :class:`TokenIdentity`.
1. **DB** (``Token`` table from Phase 4.1) the primary source. Rows
carry Argon2id hashes; the admin UI (Phase 4.3) will be the only
writer at steady state.
2. **`BOOTSTRAP_TOKENS`** env a nameplaintext map kept around for
first-run, disaster-recovery, and ``examples/`` smoke tests. These
entries live alongside DB rows in the cache and are never persisted.
Hot path is in-memory: at :meth:`start` we pull every non-revoked DB
row and stash it in a list; subsequent :meth:`verify` calls re-load
when the cache is older than ``ttl_seconds``. ``last_used_at`` updates
are coalesced into a small dict and flushed by a background task every
``flush_interval`` seconds one transaction per flush rather than one
per request.
We can't index DB rows by a derived plaintext key because Argon2 salts
are random so verify does a linear scan over candidates, calling
``argon2.PasswordHasher.verify`` on each. N is small by design (single
operator, ~10 tokens at most); the cost is irrelevant. The scan runs
through ``asyncio.to_thread`` to keep the event loop free of the ~50ms
KDF block.
The module knows nothing about HTTP frameworks. It takes a raw token
(or a verbatim ``Authorization`` header value) and returns a
:class:`TokenIdentity` (name + scope + db-id), or ``None`` for a miss.
Frontends own the 401 response shape.
""" """
from __future__ import annotations from __future__ import annotations
@@ -61,7 +36,7 @@ class TokenStoreError(ValueError):
VALID_SCOPES: frozenset[str] = frozenset({"*", "messages", "mcp", "admin", "api"}) VALID_SCOPES: frozenset[str] = frozenset({"*", "messages", "mcp", "admin", "api"})
"""The scopes a ``Token.scope`` may hold (Phase 4.3 admin UI enforces). """The scopes a ``Token.scope`` may hold.
* ``*`` wildcard, may use any frontend * ``*`` wildcard, may use any frontend
* ``messages`` Anthropic Messages frontend only * ``messages`` Anthropic Messages frontend only
@@ -69,7 +44,7 @@ VALID_SCOPES: frozenset[str] = frozenset({"*", "messages", "mcp", "admin", "api"
* ``api`` conversations API (``/api``) and its SSE * ``api`` conversations API (``/api``) and its SSE
* ``admin`` reserved for programmatic admin access; the AdminFrontend * ``admin`` reserved for programmatic admin access; the AdminFrontend
itself authenticates via session cookies, not bearer tokens, so this itself authenticates via session cookies, not bearer tokens, so this
scope is unused today and kept for forward compatibility. scope is unused today.
""" """
@@ -108,33 +83,23 @@ class _CachedToken:
hashed_value: str hashed_value: str
# Default Argon2id parameters from ``argon2-cffi`` are fine for our scope.
# They target ~50ms on a modern CPU — enough to make a stolen-hash brute
# force expensive, cheap enough to verify a handful per request.
_HASHER = PasswordHasher() _HASHER = PasswordHasher()
def hash_token(plaintext: str) -> str: def hash_token(plaintext: str) -> str:
"""Return an Argon2id hash for ``plaintext`` (admin / seed-only path). """Return an Argon2id hash for ``plaintext``, for admin/seed use."""
Phase 4.3 will call this when the admin creates a token; Phase 4.2
exposes it so smoke scripts can seed the DB without re-implementing
the same line.
"""
return _HASHER.hash(plaintext) return _HASHER.hash(plaintext)
class TokenStore: class TokenStore:
"""DB-backed verifier with in-memory cache + TTL + batched touches. """DB-backed verifier with an in-memory cache, TTL, and batched touches.
Construct in ``cli.main`` after :class:`Database` is up, then Call ``await store.start()`` after construction to prime the cache
``await store.start()`` to prime the cache and spin up the flusher and start the flusher; ``await store.stop()`` drains the touch queue.
task. ``await store.stop()`` on shutdown drains the touch queue.
Bootstrap entries (from ``BOOTSTRAP_TOKENS``) sit alongside DB rows Bootstrap entries (``BOOTSTRAP_TOKENS``) are checked first and stay
in the same lookup path; we check them first, in constant time, so usable even if the DB is unreachable; they never get flushed to
they remain usable even if the DB is unreachable. They never appear ``last_used_at`` since they have no DB row.
in ``last_used_at`` flushes because they have no DB row.
""" """
__slots__ = ( __slots__ = (
@@ -159,8 +124,11 @@ class TokenStore:
ttl_seconds: float = 30.0, ttl_seconds: float = 30.0,
flush_interval: float = 5.0, flush_interval: float = 5.0,
) -> None: ) -> None:
# Bootstrap is keyed by value internally so verify is O(1) over """Index bootstrap tokens for O(1) verify lookup.
# plaintext. Each value also keeps its name for audit lines.
Each one's scope defaults to ``"*"`` unless narrowed via
``name:value:scope``.
"""
by_value: dict[str, str] = {} by_value: dict[str, str] = {}
for name, value in (bootstrap or {}).items(): for name, value in (bootstrap or {}).items():
if not name or not value: if not name or not value:
@@ -174,9 +142,6 @@ class TokenStore:
raise TokenStoreError(msg) raise TokenStoreError(msg)
by_value[value] = name by_value[value] = name
self._bootstrap_by_value: dict[str, str] = by_value self._bootstrap_by_value: dict[str, str] = by_value
# Bootstrap entries are ``"*"`` unless the env narrows them
# (``name:value:scope``) - a webhook token in a URL should not be
# an admin token.
self._bootstrap_scopes: dict[str, str] = { self._bootstrap_scopes: dict[str, str] = {
name: (bootstrap_scopes or {}).get(name, _BOOTSTRAP_SCOPE) name: (bootstrap_scopes or {}).get(name, _BOOTSTRAP_SCOPE)
for name in by_value.values() for name in by_value.values()
@@ -192,8 +157,6 @@ class TokenStore:
self._touch_queue: dict[int, datetime] = {} self._touch_queue: dict[int, datetime] = {}
self._flusher_task: asyncio.Task[None] | None = None self._flusher_task: asyncio.Task[None] | None = None
# ---- bootstrap parsing (kept for `cli` / tests) ---------------------
@staticmethod @staticmethod
def parse_bootstrap(raw: str) -> dict[str, str]: def parse_bootstrap(raw: str) -> dict[str, str]:
"""Parse ``name1:value1,name2:value2[:scope]`` (``BOOTSTRAP_TOKENS``).""" """Parse ``name1:value1,name2:value2[:scope]`` (``BOOTSTRAP_TOKENS``)."""
@@ -210,15 +173,9 @@ class TokenStore:
@classmethod @classmethod
def from_env(cls, raw: str, db: Database | None = None) -> TokenStore: def from_env(cls, raw: str, db: Database | None = None) -> TokenStore:
"""Legacy entrypoint: bootstrap-only (or bootstrap + db). """Bootstrap-only (or bootstrap + db) convenience constructor."""
Phase 1.3 call sites still expect a one-liner; we keep the
classmethod so they don't have to learn the new constructor.
"""
return cls(db, bootstrap=cls.parse_bootstrap(raw)) return cls(db, bootstrap=cls.parse_bootstrap(raw))
# ---- lifecycle ------------------------------------------------------
async def start(self) -> None: async def start(self) -> None:
"""Prime the cache and (if a DB is attached) start the flusher loop.""" """Prime the cache and (if a DB is attached) start the flusher loop."""
await self._refresh() await self._refresh()
@@ -242,19 +199,19 @@ class TokenStore:
self._bootstrap_scopes[name] = scope self._bootstrap_scopes[name] = scope
async def invalidate(self) -> None: async def invalidate(self) -> None:
"""Force the next verify to re-read from DB (Phase 4.3 admin hook).""" """Force the next verify to re-read from DB."""
self._loaded_at = 0.0 self._loaded_at = 0.0
# ---- verify path ----------------------------------------------------
async def verify(self, token: str | None) -> TokenIdentity | None: async def verify(self, token: str | None) -> TokenIdentity | None:
"""Return the matching identity, or ``None`` for unknown/empty tokens.""" """Return the matching identity, or ``None`` for unknown/empty tokens.
Bootstrap entries are checked first via constant-time compare;
the DB cache list is swapped rather than mutated on refresh, so
a concurrent reload can't affect an in-progress scan.
"""
if not token: if not token:
return None return None
# Bootstrap first: constant-time compare per entry, never hits DB.
# `compare_digest` is overkill for a name→value lookup but cheap
# and removes one timing variable for free.
for value, name in self._bootstrap_by_value.items(): for value, name in self._bootstrap_by_value.items():
if hmac.compare_digest(token, value): if hmac.compare_digest(token, value):
return TokenIdentity( return TokenIdentity(
@@ -268,9 +225,6 @@ class TokenStore:
await self._ensure_fresh() await self._ensure_fresh()
# Snapshot the cache reference so a refresh mid-scan doesn't
# surprise us. List itself is immutable per refresh (we swap,
# not mutate).
cache = self._cache cache = self._cache
for entry in cache: for entry in cache:
try: try:
@@ -306,17 +260,18 @@ class TokenStore:
def __bool__(self) -> bool: def __bool__(self) -> bool:
return bool(self._cache) or bool(self._bootstrap_by_value) return bool(self._cache) or bool(self._bootstrap_by_value)
# ---- internals ------------------------------------------------------
async def _ensure_fresh(self) -> None: async def _ensure_fresh(self) -> None:
"""Reload the cache when stale.
Double-checked under the lock so concurrent callers don't all
trigger a reload.
"""
if self._db is None: if self._db is None:
return return
now = time.monotonic() now = time.monotonic()
if now - self._loaded_at <= self._ttl: if now - self._loaded_at <= self._ttl:
return return
async with self._lock: async with self._lock:
# Re-check under the lock — first arrival reloaded, others
# should fall through.
now = time.monotonic() now = time.monotonic()
if now - self._loaded_at <= self._ttl: if now - self._loaded_at <= self._ttl:
return return
@@ -331,8 +286,6 @@ class TokenStore:
next_cache: list[_CachedToken] = [] next_cache: list[_CachedToken] = []
for row in rows: for row in rows:
if row.id is None: if row.id is None:
# Defensive: SQLModel will assign an id on insert; a
# None here would mean someone handed us an unsaved row.
continue continue
next_cache.append( next_cache.append(
_CachedToken( _CachedToken(
@@ -353,20 +306,20 @@ class TokenStore:
await self._flush_now() await self._flush_now()
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
except Exception: # noqa: BLE001 — never let the flusher die silently except Exception: # noqa: BLE001
_log.exception("token flusher crashed; touches will stop") _log.exception("token flusher crashed; touches will stop")
async def _flush_now(self) -> None: async def _flush_now(self) -> None:
"""Flush queued touches in one transaction.
Detaches the queue first so concurrent ``verify()`` writes land
in the next flush.
"""
if self._db is None or not self._touch_queue: if self._db is None or not self._touch_queue:
return return
# Detach the queue so concurrent verify() writes don't bleed
# into the in-flight transaction.
pending, self._touch_queue = self._touch_queue, {} pending, self._touch_queue = self._touch_queue, {}
async with self._db.session() as session: async with self._db.session() as session:
for token_id in pending: for token_id in pending:
# We don't pass the timestamp through — `touch_token`
# stamps `now` itself, and we'd rather have one source
# of truth than reconcile clocks.
await touch_token(session, token_id=token_id) await touch_token(session, token_id=token_id)
_log.debug("flushed %d token touch(es)", len(pending)) _log.debug("flushed %d token touch(es)", len(pending))
+16 -67
View File
@@ -1,33 +1,7 @@
"""Recognise credentials in free text and mask them. """Recognise credentials in free text and mask them.
``docker logs`` is not a vault: the stack's stdout lands in an unrotated Used to filter log output and to redact what MCP tools hand back to
json file on the host, and the model itself can page through it (the the model (see :mod:`beaver_gateway.mcp.redacting`).
komodo tool has ``logs`` / ``search_logs``). Nobody has to write a
careless log call for a secret to end up there two habits do it on
their own:
* ``httpx`` logs ``HTTP Request: POST <full url>`` at ``INFO``, so every
request to an upstream that keeps its credential *in* the URL prints
the credential once per call;
* a transport error carries that same URL through the traceback, which
``_log.exception`` writes out in full.
Two of our upstreams are exactly that shape: a Google Calendar
``.../private-<token>/basic.ics`` feed and a USOS ``?key=<token>`` ical
feed, both handed to the calendar MCP as query parameters.
So: silence the HTTP clients' per-request chatter (nothing here reads
it), and run one redaction pass over every formatted record as a second
line of defence.
The same pass guards the wider channel what an MCP tool hands back to
the model (see :mod:`beaver_gateway.mcp.redacting`). :func:`redact`
recognises three things: literal values of the secret-looking
environment variables this process was started with, credentials that
live in URLs and headers, and ``NAME: value`` assignments whose *name*
says the value is a credential. The last one is what covers a config
dump of a stack whose secrets this process never held and it keeps
the name, masking only the value, so the dump still says what is set.
""" """
from __future__ import annotations from __future__ import annotations
@@ -43,39 +17,29 @@ if TYPE_CHECKING:
MASK = "<…>" MASK = "<…>"
# Loggers that print request URLs at INFO. We keep their warnings.
CHATTY: tuple[str, ...] = ("httpx", "httpcore", "aiohttp.client", "urllib3") CHATTY: tuple[str, ...] = ("httpx", "httpcore", "aiohttp.client", "urllib3")
"""Loggers that print request URLs at ``INFO``; only their warnings are kept."""
# Loggers that bring their own handlers (uvicorn re-runs ``dictConfig``
# when a server starts), so wrapping the root formatter misses them.
# ``dictConfig`` drops a logger's handlers but keeps its filters.
OWN_HANDLERS: tuple[str, ...] = ("uvicorn", "uvicorn.access", "uvicorn.error") OWN_HANDLERS: tuple[str, ...] = ("uvicorn", "uvicorn.access", "uvicorn.error")
"""Loggers whose own handlers survive ``dictConfig``, unreached by the
root formatter."""
# Names whose *value* is a credential — matched as a substring, so
# ``FIREFLY_PAT`` and ``T3_MAC_TOKEN`` both qualify. Used twice: to pick
# which environment variables contribute literal values, and to decide
# whether a ``NAME: value`` line in a config dump should keep its value.
# ``MCP`` is in here because every MCP URL in this stack carries its
# credential in the query string.
_SECRET_NAME = re.compile( _SECRET_NAME = re.compile(
r"TOKEN|SECRET|KEY|PASSWORD|PASS\b|PAT\b|BEARER|CREDENTIAL" r"TOKEN|SECRET|KEY|PASSWORD|PASS\b|PAT\b|BEARER|CREDENTIAL"
r"|MCPS?\b|AUTH|PRIVATE|SESSION|DSN", r"|MCPS?\b|AUTH|PRIVATE|SESSION|DSN",
re.IGNORECASE, re.IGNORECASE,
) )
"""Substring match for names whose *value* is a credential."""
_MIN_SECRET = 8 _MIN_SECRET = 8
# A query parameter may sit behind a plain ``?``/``&`` or behind their
# percent-encoded twins when a whole URL is nested in another one.
_LEFT = r"(?:(?<![A-Za-z0-9])|(?<=%26)|(?<=%3F))" _LEFT = r"(?:(?<![A-Za-z0-9])|(?<=%26)|(?<=%3F))"
# ...and runs until the next separator, encoded ``&`` included.
_VALUE = r"(?:(?!%26)[^&\s\"'<>,;)\]}])+" _VALUE = r"(?:(?!%26)[^&\s\"'<>,;)\]}])+"
"""Query-parameter boundaries: a value may follow a plain or
percent-encoded ``?``/``&`` and runs to the next separator."""
# Names worth masking where they sit in a dump next to their password,
# but not worth hunting for as literals across every log line: a
# username is short and ordinary, and ``CLAUDE_RUNNER_USER=beaver-runner``
# turned into a global literal would mask half the entrypoint script.
_SECRET_IN_DUMP = re.compile(r"USER\b|LOGIN\b", re.IGNORECASE) _SECRET_IN_DUMP = re.compile(r"USER\b|LOGIN\b", re.IGNORECASE)
"""Masked only when paired with a password in a dump, not hunted as a global literal."""
def is_secret_name(name: str) -> bool: def is_secret_name(name: str) -> bool:
@@ -102,13 +66,10 @@ def _mask_named_value(match: re.Match[str]) -> str:
type _Repl = str | Callable[[re.Match[str]], str] type _Repl = str | Callable[[re.Match[str]], str]
_RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = ( _RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = (
# Google Calendar's secret path segment; the ``/basic.ics`` after it
# survives, encoded or not, so the line still says what it fetched.
( (
re.compile(r"(private-)(?:(?!%2F)[A-Za-z0-9_%-]){8,}", re.IGNORECASE), re.compile(r"(private-)(?:(?!%2F)[A-Za-z0-9_%-]){8,}", re.IGNORECASE),
r"\1" + MASK, r"\1" + MASK,
), ),
# ``key=…``, ``api_key=…``, ``token=…``, … in a query string.
( (
re.compile( re.compile(
_LEFT _LEFT
@@ -120,22 +81,13 @@ _RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = (
), ),
r"\1\2" + MASK, r"\1\2" + MASK,
), ),
# ``Authorization: Bearer …`` in a header dump; scheme goes too.
( (
re.compile( re.compile(
r"(authorization[\"']?\s*[:=]\s*[\"']?)(?:\S+\s+)?\S+", re.IGNORECASE r"(authorization[\"']?\s*[:=]\s*[\"']?)(?:\S+\s+)?\S+", re.IGNORECASE
), ),
r"\1" + MASK, r"\1" + MASK,
), ),
# A bare ``Bearer <token>`` anywhere else.
(re.compile(r"\b(bearer)\s+[A-Za-z0-9._~+/=-]{8,}", re.IGNORECASE), r"\1 " + MASK), (re.compile(r"\b(bearer)\s+[A-Za-z0-9._~+/=-]{8,}", re.IGNORECASE), r"\1 " + MASK),
# ``NAME: value`` / ``NAME=value`` at the start of a line — the shape
# of a compose ``environment:`` block, an ``.env`` file, a printed
# settings object. Anchored to the line so that an ``actor=token:foo``
# in the middle of a log line doesn't swallow the rest of it, and
# ``=`` must be tight (``NAME=value``, the way env files write it) so
# that ``SESSION_SECRET = os.environ[…]`` in source the dispatcher is
# reading through t3code stays readable.
( (
re.compile( re.compile(
r"""(?m)^(?P<lead>[ \t]*(?:-[ \t]+)?["']?) r"""(?m)^(?P<lead>[ \t]*(?:-[ \t]+)?["']?)
@@ -146,9 +98,6 @@ _RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = (
), ),
_mask_named_value, _mask_named_value,
), ),
# The same at the start of a line in lower case, but only for names
# that are unambiguously a credential — ``_SECRET_NAME`` is too broad
# here, it would read the ``mcp:`` prefix of a log line as one.
( (
re.compile( re.compile(
r"""(?m)^(?P<lead>[ \t]*(?:-[ \t]+)?["']?) r"""(?m)^(?P<lead>[ \t]*(?:-[ \t]+)?["']?)
@@ -160,12 +109,7 @@ _RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = (
), ),
_mask_named_value, _mask_named_value,
), ),
# ``scheme://user:password@host`` — the credential a connection
# string carries. Host and database stay, so the line still locates
# the service it failed to reach.
(re.compile(r"(://[^/\s:@]+:)[^@\s/]+(@)"), r"\1" + MASK + r"\2"), (re.compile(r"(://[^/\s:@]+:)[^@\s/]+(@)"), r"\1" + MASK + r"\2"),
# ``"name": "value"`` anywhere — the same idea for one-line JSON,
# where the quotes make the end of the value unambiguous.
( (
re.compile( re.compile(
r'(?P<lead>")(?P<name>[^"\n]{1,64})(?P<sep>"\s*:\s*")' r'(?P<lead>")(?P<name>[^"\n]{1,64})(?P<sep>"\s*:\s*")'
@@ -174,6 +118,10 @@ _RULES: tuple[tuple[re.Pattern[str], _Repl], ...] = (
_mask_named_value, _mask_named_value,
), ),
) )
"""Redaction rules applied in order: calendar tokens, query-string secrets,
``Authorization`` headers, bare bearer tokens, ``NAME: value``/``NAME=value``
lines (upper- then lower-case names), connection-string passwords, and
one-line JSON ``"name": "value"`` pairs."""
_ENV_SECRETS: list[str] = [] _ENV_SECRETS: list[str] = []
"""Literal secret values to mask, longest first. Filled by :func:`install`.""" """Literal secret values to mask, longest first. Filled by :func:`install`."""
@@ -255,9 +203,10 @@ class RedactFilter(logging.Filter):
""" """
def filter(self, record: logging.LogRecord) -> bool: def filter(self, record: logging.LogRecord) -> bool:
"""Let a record through unmasked if its template is malformed."""
try: try:
message = record.getMessage() message = record.getMessage()
except (TypeError, ValueError): # a broken template is not ours to fix except (TypeError, ValueError):
return True return True
masked = redact(message) masked = redact(message)
if masked != message: if masked != message:
+23 -17
View File
@@ -1,8 +1,8 @@
"""Process-wide configuration loaded from environment (`.env`). """Process-wide configuration loaded from environment (`.env`).
Everything that varies between deployments (creds, paths, ports) lives here. Everything that varies between deployments (creds, paths, ports) lives
User-facing agent/MCP/frontend definitions live in ``/config/config.py`` here. User-facing agent/MCP/frontend definitions live in ``config_path``
and are loaded by ``config_loader`` (Phase 0.3). and are loaded by :mod:`beaver_gateway.config`.
""" """
from __future__ import annotations from __future__ import annotations
@@ -20,12 +20,22 @@ class Settings(BaseSettings):
) )
database_url: str database_url: str
"""SQLAlchemy async URL for the gateway's database."""
admin_user: str admin_user: str
"""Username for the admin UI login."""
admin_pass: str admin_pass: str
"""Password for the admin UI login."""
session_secret: str session_secret: str
"""Secret used to sign admin UI session cookies."""
internal_mcp_port: int = 8765 internal_mcp_port: int = 8765
"""Port the gateway's own internal MCP server listens on."""
config_path: Path = Path("/config/config.py") config_path: Path = Path("/config/config.py")
"""Where the user-facing agent/MCP/frontend config module lives."""
claude_runner_user: str | None = None claude_runner_user: str | None = None
"""Unix user the claude subprocess runs as. Needs the gateway to be root.""" """Unix user the claude subprocess runs as. Needs the gateway to be root."""
@@ -34,23 +44,19 @@ class Settings(BaseSettings):
"""``HOME`` for the claude subprocess (its ``~/.claude`` lives there).""" """``HOME`` for the claude subprocess (its ``~/.claude`` lives there)."""
raycast_bearer: str | None = None raycast_bearer: str | None = None
"""Bearer token the gateway presents to the Raycast API."""
raycast_config_path: Path = Path("/config/raycast.json") raycast_config_path: Path = Path("/config/raycast.json")
"""Where cached Raycast API config (fetched on first use) is stored."""
raycast_device_id: str | None = None raycast_device_id: str | None = None
"""64-hex-char stable per-install id. Required when any ``RaycastAgent`` """Stable per-install id (``secrets.token_hex(32)``); required when any
is configured generate once via ``secrets.token_hex(32)`` and keep ``RaycastAgent`` is configured, so restarts don't look like a new device."""
in ``.env`` so Raycast doesn't see every restart as a new device."""
raycast_locale: str = "en-US" raycast_locale: str = "en-US"
"""``Accept-Language`` header sent by the shared ``raycast_api.Client`` """``Accept-Language`` sent to the Raycast API and the default locale for
and the default locale used to render auto ``UserPreferences``. One auto ``UserPreferences``. One value per gateway (one shared Client)."""
value per gateway since we open one Client total."""
bootstrap_tokens: str = "" bootstrap_tokens: str = ""
"""Out-of-band token seed: ``name1:value1,name2:value2``. Empty is """Out-of-band token seed: ``name1:value1,name2:value2``. Layers
fine DB-issued tokens (admin UI) carry steady-state auth. alongside DB-issued tokens; used for first-run setup and recovery."""
This env channel layers alongside the DB store and is kept for
first-run setup, disaster recovery (admin password lost), and
``examples/`` smoke tests. Bootstrap tokens implicitly carry
scope ``"*"`` and never get a ``last_used_at`` stamp (no DB row).
"""
+2 -8
View File
@@ -1,13 +1,7 @@
"""Rewrite ``usage.cost_usd``/``model_usage`` as per-turn deltas. """Rewrite ``usage.cost_usd``/``model_usage`` as per-turn deltas.
Until 2026-09-01 the gateway stored ``ResultMessage.total_cost_usd`` verbatim, Older rows stored a cumulative total, which overstates daily sums. Safe
which is cumulative for the claude process - summing the column overstated a to rerun; reads ``DATABASE_URL`` like the gateway.
day by an order of magnitude. Run once after deploying the delta-aware
:func:`beaver_gateway.storage.append_usage`::
python -m beaver_gateway.storage.backfill_usage
Reads ``DATABASE_URL`` like the gateway (``.env`` included); safe to rerun.
""" """
from __future__ import annotations from __future__ import annotations
+8 -24
View File
@@ -1,15 +1,8 @@
"""Async ``Database`` wrapper + the bare-minimum CRUD helpers. """Async ``Database`` wrapper + the bare-minimum CRUD helpers.
Async to match the rest of the stack (aiohttp, uvicorn, claude-code-api). psycopg3 gives native async support for Postgres; SQLite goes through
psycopg3 has native async support ``postgresql+psycopg://...`` works ``aiosqlite``. Helpers take an ``AsyncSession`` so callers can batch
with ``create_async_engine`` directly. SQLite goes through ``aiosqlite`` operations into one transaction.
(``sqlite+aiosqlite://...``); user-facing config still uses the plain
``sqlite:///`` form and we normalise the URL here, so nothing leaks into
``.env`` / docker-compose.
No repository layer (PLAN §4.1 explicitly waives it). Helpers take an
``AsyncSession`` so callers can batch operations into one transaction
(e.g. touch ``last_used_at`` + write an audit line on the same request).
""" """
from __future__ import annotations from __future__ import annotations
@@ -130,11 +123,8 @@ def _column_default(column: Any) -> str | None:
return None return None
# ---- Token CRUD ---------------------------------------------------------
async def list_active_tokens(session: AsyncSession) -> Sequence[Token]: async def list_active_tokens(session: AsyncSession) -> Sequence[Token]:
"""Return every non-revoked token (Phase 4.2 seeds the cache from this).""" """Return every non-revoked token."""
stmt = select(Token).where(Token.revoked_at.is_(None)) # ty: ignore[unresolved-attribute] stmt = select(Token).where(Token.revoked_at.is_(None)) # ty: ignore[unresolved-attribute]
result = await session.exec(stmt) result = await session.exec(stmt)
return result.all() return result.all()
@@ -143,7 +133,7 @@ async def list_active_tokens(session: AsyncSession) -> Sequence[Token]:
async def list_tokens( async def list_tokens(
session: AsyncSession, *, include_revoked: bool = False session: AsyncSession, *, include_revoked: bool = False
) -> Sequence[Token]: ) -> Sequence[Token]:
"""Return tokens ordered newest-first (Phase 4.3 admin table).""" """Return tokens ordered newest-first."""
stmt = select(Token).order_by(Token.created_at.desc()) # ty: ignore[unresolved-attribute] stmt = select(Token).order_by(Token.created_at.desc()) # ty: ignore[unresolved-attribute]
if not include_revoked: if not include_revoked:
stmt = stmt.where(Token.revoked_at.is_(None)) # ty: ignore[unresolved-attribute] stmt = stmt.where(Token.revoked_at.is_(None)) # ty: ignore[unresolved-attribute]
@@ -174,7 +164,7 @@ async def revoke_token(session: AsyncSession, *, token_id: int) -> bool:
async def touch_token(session: AsyncSession, *, token_id: int) -> None: async def touch_token(session: AsyncSession, *, token_id: int) -> None:
"""Bump ``last_used_at``. Phase 4.2 batches these — not per-request.""" """Bump ``last_used_at`` for one token."""
row = await session.get(Token, token_id) row = await session.get(Token, token_id)
if row is None: if row is None:
return return
@@ -183,9 +173,6 @@ async def touch_token(session: AsyncSession, *, token_id: int) -> None:
await session.commit() await session.commit()
# ---- Audit --------------------------------------------------------------
async def append_audit( async def append_audit(
session: AsyncSession, session: AsyncSession,
*, *,
@@ -223,9 +210,6 @@ async def list_audit_records(
return result.all() return result.all()
# ---- Usage --------------------------------------------------------------
async def append_usage(session: AsyncSession, row: Usage) -> None: async def append_usage(session: AsyncSession, row: Usage) -> None:
"""Persist a turn's usage with ``cost_usd`` / ``model_usage`` as per-turn deltas. """Persist a turn's usage with ``cost_usd`` / ``model_usage`` as per-turn deltas.
@@ -252,8 +236,6 @@ async def _last_usage(session: AsyncSession, session_id: str | None) -> Usage |
return (await session.exec(stmt)).first() return (await session.exec(stmt)).first()
# Counters inside ``ResultMessage.model_usage[model]``; everything else there
# (provider, costBasis, contextWindow, maxOutputTokens, ...) is a constant.
_MODEL_USAGE_COUNTERS = ( _MODEL_USAGE_COUNTERS = (
"inputTokens", "inputTokens",
"outputTokens", "outputTokens",
@@ -262,6 +244,8 @@ _MODEL_USAGE_COUNTERS = (
"webSearchRequests", "webSearchRequests",
"costUSD", "costUSD",
) )
"""Counters inside ``ResultMessage.model_usage[model]``; other keys there
are constants."""
def usage_deltas(row: Usage, prev: Usage | None) -> None: def usage_deltas(row: Usage, prev: Usage | None) -> None:
+15 -32
View File
@@ -1,23 +1,8 @@
"""SQLModel tables. """SQLModel tables.
Flat tables, no FK relationships modelled (``actor`` and Flat, no FK relationships modelled ``actor``/``agent_name`` are plain
``agent_name`` are stored as strings joining audittoken by name is strings. Datetimes are stored UTC via ``default_factory`` so SQLite and
fine at this volume; we'll introduce FKs when the admin UI actually Postgres behave identically.
demands them).
The ``Conversation`` + ``ConversationMessage`` pair persists chat
history per frontend so we can survive cache misses without losing
tool-call memory. The gateway is now stateful about conversation
content (we keep the raw Anthropic-shape message list including
``tool_use`` / ``tool_result`` blocks); the live ``claude-code-api``
session pool stays the source of truth for *fingerprints*, and the DB
mirrors what we'd want to re-seed if a session evicts. See
``core/conversation_store.py`` for the diff-and-fork logic and
``frontends/markdown/frontend.py`` for the integration point.
Datetimes are stored UTC; we set ``default_factory`` rather than relying
on DB defaults so SQLite + Postgres behave identically. Every row that
needs an id uses ``Optional[int]`` so SQLAlchemy can autoincrement.
""" """
from __future__ import annotations from __future__ import annotations
@@ -37,10 +22,8 @@ def _utcnow() -> datetime:
class Token(SQLModel, table=True): class Token(SQLModel, table=True):
"""Bearer token issued to an external caller. """Bearer token issued to an external caller.
``hashed_value`` holds the Argon2 hash (Phase 4.2 until then, ``hashed_value`` holds the Argon2 hash. Plaintext is shown to the
rows are written by tests / the admin UI, not by ``TokenStore``). user **once** at creation and then discarded.
Plaintext is shown to the user **once** at creation and then
discarded.
""" """
id: int | None = Field(default=None, primary_key=True) id: int | None = Field(default=None, primary_key=True)
@@ -74,7 +57,7 @@ class AuditLog(SQLModel, table=True):
class Conversation(SQLModel, table=True): class Conversation(SQLModel, table=True):
"""One conversation (§3.1): a master thread, a branch, a deep chat or a job. """One conversation: a master thread, a branch, a deep chat or a job.
``external_id`` is the public id (uuid) every frontend, the API and the ``external_id`` is the public id (uuid) every frontend, the API and the
usage table refer to; ``frontend`` names the frontend that created the usage table refer to; ``frontend`` names the frontend that created the
@@ -83,7 +66,7 @@ class Conversation(SQLModel, table=True):
``running_turn`` survives a restart: a non-null value at startup means ``running_turn`` survives a restart: a non-null value at startup means
the gateway died mid-turn and the transcript needs its open ``tool_use`` the gateway died mid-turn and the transcript needs its open ``tool_use``
closed before the session is resumed (``core/conversations``). closed before the session is resumed.
""" """
__tablename__ = "conversations" __tablename__ = "conversations"
@@ -119,10 +102,10 @@ class Conversation(SQLModel, table=True):
class ConversationBinding(SQLModel, table=True): class ConversationBinding(SQLModel, table=True):
"""Where a conversation shows up: a Telegram topic id, a vault-relative path. """Where a conversation shows up: a Telegram topic id, a vault-relative path.
Invariant (§3.1): at most one *visible* binding per frontend per Invariant: at most one *visible* binding per frontend per conversation -
conversation - enforced by the partial unique index. The same external enforced by the partial unique index. The same external id may point at
id may point at several conversations over time (a renamed topic gets a several conversations over time (a renamed topic gets a new branch),
new branch), only one of them visible. only one of them visible.
""" """
__tablename__ = "conversation_bindings" __tablename__ = "conversation_bindings"
@@ -147,7 +130,7 @@ class ConversationBinding(SQLModel, table=True):
class InjectQueueItem(SQLModel, table=True): class InjectQueueItem(SQLModel, table=True):
"""Persisted per-conversation queue (§3.4), ``urgent > user > wake > normal``. """Persisted per-conversation queue, priority ``urgent > user > wake > normal``.
``status`` walks ``queued -> running -> done``; a row still ``running`` ``status`` walks ``queued -> running -> done``; a row still ``running``
at startup was cut by a restart and becomes ``interrupted`` - it is at startup was cut by a restart and becomes ``interrupted`` - it is
@@ -171,7 +154,7 @@ class InjectQueueItem(SQLModel, table=True):
class TelegramUpdate(SQLModel, table=True): class TelegramUpdate(SQLModel, table=True):
"""Inbox of the Telegram frontend (§3.8). """Inbox of the Telegram frontend.
The update is stored before the poll offset moves past it and handled The update is stored before the poll offset moves past it and handled
from here, so a restart neither loses nor duplicates a message. from here, so a restart neither loses nor duplicates a message.
@@ -189,7 +172,7 @@ class TelegramUpdate(SQLModel, table=True):
class Delivery(SQLModel, table=True): class Delivery(SQLModel, table=True):
"""Outbox (§3.8): a reply is a row first and a Telegram message second. """Outbox: a reply is a row first and a Telegram message second.
``status`` walks ``queued -> sent`` with retries on the way, ``failed`` ``status`` walks ``queued -> sent`` with retries on the way, ``failed``
only when Telegram rejects the row for good (unknown thread, blocked only when Telegram rejects the row for good (unknown thread, blocked
@@ -361,7 +344,7 @@ class RateLimit(SQLModel, table=True):
class JobRunRecord(SQLModel, table=True): class JobRunRecord(SQLModel, table=True):
"""One finished run of a scheduler job (§3.6): what fired it, how it ended. """One finished run of a scheduler job: what fired it, how it ended.
``error`` keeps the head of the traceback of a failed run; ``payload`` ``error`` keeps the head of the traceback of a failed run; ``payload``
is the trigger's data (``{}`` for cron) and should stay small. is the trigger's data (``{}`` for cron) and should stay small.
+2 -4
View File
@@ -1,9 +1,7 @@
"""``SessionStore`` adapter for the Claude Agent SDK on top of :class:`Database`. """``SessionStore`` adapter for the Claude Agent SDK on top of :class:`Database`.
Entries are stored verbatim as JSON (``jsonb`` on Postgres), ordered by a Entries are JSON rows ordered by a per-key ``seq``; ``append`` is
per-key ``seq``. ``append`` is idempotent on ``entry["uuid"]`` because the idempotent on ``entry["uuid"]`` since the SDK may re-deliver a batch.
SDK re-delivers a batch on retry; entries without a uuid are appended as
they come. Runs on SQLite too - the conformance suite uses that in tests.
""" """
from __future__ import annotations from __future__ import annotations
+1
View File
@@ -0,0 +1 @@
"""Watching a directory of notes for the envelope."""
+4 -7
View File
@@ -1,11 +1,8 @@
"""Vault watcher for the envelope (§3.5, §4.6). """Vault watcher for the envelope.
A content snapshot is taken at every envelope; a change is the diff of the A content snapshot is taken at every envelope; a change is the diff of a
file against that snapshot, "added lines only". ``watchfiles`` delivers file against that snapshot ("added lines only"). ``watchfiles`` debounces
paths with a settle debounce (Sync writes files in pieces), the change delivery, so a change carries the file's mtime, not when it landed.
carries the file's mtime rather than the moment it landed. Which files
show a full diff, which show up by name and which are ignored is a set of
glob patterns the config hands in - the gateway itself knows no path.
""" """
from __future__ import annotations from __future__ import annotations