feat(stream): headless claude -p transport behind a transport flag
The PTY transport stands in for a protocol that did not exist when it was written: it pastes bracketed text into claude's TUI, guesses when the Ink render loop has settled, re-presses Enter when the paste is swallowed, and tails the session JSONL at 100ms. `claude -p` with stream-json on both pipes is that protocol, so add it as a second transport and let callers pick with `BackendOptions.transport`. Measured on one host, same model and prompt, cold single-reply turn: pty burns 2.3s on TUI readiness before the prompt is even submitted (first event 2.5s, turn 5.2s); stream_json burns none (first event 0.3s, first token 1.9s, turn 2.7s). The gap is what `startup_delay`'s 60s cap exists to survive on slow hardware. Everything the old transport relies on carries over unchanged and was verified against a real CLI: multi-turn over one live process, MCP via --mcp-config, --resume, and `native_jsonl` history seeding. Two things are new rather than equal: `result` is native (so usage is the turn's aggregate and durations are real, not synthesized), and `include_partial_messages` yields token-level `StreamEvent`s — which the JSONL, holding only finished blocks, could never provide. Notably, assistant records carry `stop_reason: null` in this mode even at the end of a turn, so `result` is not just tidier than the PTY path's terminal-stop_reason heuristic, it is the only correct signal. Default stays `pty`; nothing changes for existing callers. Also: extract process-group bookkeeping into `procgroup` so both transports sweep claude's children on shutdown, and fix a stale assertion in test_backend that still expected the `is_error: None` that `_block_to_dict` deliberately stopped emitting.
This commit is contained in:
@@ -2,10 +2,27 @@
|
||||
[](https://sladge.net)
|
||||
|
||||
Python wrapper around the `claude` CLI for subscription-mode (no API key)
|
||||
backends. Drives one long-running interactive `claude` per conversation via
|
||||
a PTY and reads events from the JSONL session file; the public surface is
|
||||
Anthropic-Messages-API shaped so a gateway in front of it is a one-liner
|
||||
serializer away.
|
||||
backends. Drives one long-running `claude` per conversation and yields
|
||||
typed events; the public surface is Anthropic-Messages-API shaped so a
|
||||
gateway in front of it is a one-liner serializer away.
|
||||
|
||||
Two transports, picked with `BackendOptions.transport`:
|
||||
|
||||
- **`stream_json`** — `claude -p` with stream-json on stdin and stdout.
|
||||
Prompts go into a pipe, events come back as they happen, and a native
|
||||
`result` record closes each turn. Prefer this.
|
||||
- **`pty`** (default, for backwards compatibility) — drives the
|
||||
interactive TUI through a pseudo-tty and tails the session JSONL. It
|
||||
has to guess when the TUI is ready to accept a paste, re-press Enter
|
||||
when the paste is swallowed, and poll a file for output. Keep it only
|
||||
if you need something the TUI alone has (e.g. `--remote-control`).
|
||||
|
||||
Measured on the same host, same model, same prompt — cold turn, single
|
||||
reply: `pty` spends 2.3s on TUI readiness before the prompt is even
|
||||
submitted (first event 2.5s, turn 5.2s); `stream_json` spends none
|
||||
(first event 0.3s, first token 1.9s, turn 2.7s). The gap widens on slow
|
||||
hardware, which is what the PTY transport's 60s `startup_delay` cap
|
||||
exists to survive.
|
||||
|
||||
Not affiliated with Anthropic. You need a working subscription, the
|
||||
`claude` CLI on PATH, and to have run `claude /login` once.
|
||||
@@ -27,7 +44,12 @@ import asyncio
|
||||
from claude_code_api import BackendOptions, ClaudeCodeBackend
|
||||
|
||||
async def main() -> None:
|
||||
opts = BackendOptions(cwd="/path/to/project", dangerously_skip_permissions=True)
|
||||
opts = BackendOptions(
|
||||
cwd="/path/to/project",
|
||||
dangerously_skip_permissions=True,
|
||||
transport="stream_json",
|
||||
include_partial_messages=True, # token-level deltas, stream_json only
|
||||
)
|
||||
async with ClaudeCodeBackend(opts) as backend:
|
||||
async for event in backend.complete(
|
||||
[{"role": "user", "content": "say hi"}]
|
||||
@@ -72,19 +94,38 @@ session orchestration.
|
||||
|
||||
## How a turn works
|
||||
|
||||
Steps 1–3 are shared by both transports:
|
||||
|
||||
1. The backend looks up a live session by `hash_history(messages[:-1])`.
|
||||
If one matches, the new user message goes straight into its PTY.
|
||||
If one matches, the new user message goes straight to it.
|
||||
2. If nothing matches and `messages[:-1]` is empty, a fresh `claude` is
|
||||
spawned with a brand-new `--session-id`.
|
||||
3. If `messages[:-1]` is non-empty (a continuation we don't have a live
|
||||
PTY for — e.g. after restart), the backend writes a hand-crafted
|
||||
process for — e.g. after restart), the backend writes a hand-crafted
|
||||
JSONL transcript at `~/.claude/projects/<key>/<id>.jsonl` and spawns
|
||||
`claude --resume <id>`. That is the `native_jsonl` injection mode;
|
||||
the fallback is `concat_message`, which folds the prior history into
|
||||
one large first prompt.
|
||||
|
||||
Then, on `stream_json`:
|
||||
|
||||
4. The prompt is written to stdin as one JSON line. No readiness wait —
|
||||
the CLI buffers stdin, so this works even before it has finished
|
||||
booting.
|
||||
5. stdout is read line by line and each record normalized into a typed
|
||||
`Event`. With `include_partial_messages` the raw Anthropic
|
||||
`message_stream` events also come through as `StreamEvent`, with
|
||||
`index` rebased across the turn's several API requests.
|
||||
6. The turn closes on the native `result` record, which supplies the
|
||||
`ResultMessage` verbatim — aggregate usage, real durations, cost.
|
||||
|
||||
Or, on `pty`:
|
||||
|
||||
4. The PTY's stdout is drained continuously by a background thread; we
|
||||
never read events from there. The JSONL file is tailed at 100ms
|
||||
cadence and each new record is normalized into a typed `Event`.
|
||||
never read events from there. The prompt is pasted in
|
||||
bracketed-paste framing followed by a separate Enter, and re-sent if
|
||||
the JSONL shows no sign it registered. The JSONL file is tailed at
|
||||
100ms cadence and each new record is normalized into a typed `Event`.
|
||||
5. The turn closes on the first `assistant` record with `stop_reason ∈
|
||||
{end_turn, max_tokens, stop_sequence, refusal}`. A `ResultMessage`
|
||||
is synthesized from its `usage` and yielded last.
|
||||
@@ -100,7 +141,7 @@ session orchestration.
|
||||
|
||||
```bash
|
||||
uv run pytest # unit tests (fast, no real claude)
|
||||
RUN_CLAUDE_SMOKE=1 uv run pytest tests/test_pty.py tests/test_turn.py tests/test_backend.py
|
||||
RUN_CLAUDE_SMOKE=1 uv run pytest tests/test_pty.py tests/test_turn.py tests/test_backend.py tests/test_stream.py
|
||||
```
|
||||
|
||||
The smoke-marked tests spawn a real `claude` process and need a logged-in
|
||||
|
||||
Reference in New Issue
Block a user