Files
bruno 7b5aa6fcc9
CI / fmt, clippy and tests (ubuntu-latest) (push) Successful in 21m0s
CI / browser end-to-end (Chromium + fixture) (push) Skipped
CI / fmt, clippy and tests (windows-latest) (push) Canceled after 0s
Add initial llm-bridge project
2026-09-18 18:01:07 -04:00

21 KiB

llm-bridge — llm-gateway

A local, OpenAI-compatible HTTP API that answers by driving a real, signed-in browser session on an LLM web UI instead of a paid API key.

OpenAI SDK / curl / Continue / Cursor
                |
                |  POST /v1/chat/completions   (OpenAI format)
                v
        llm-gateway (Axum, Rust)
                |
                |  Chrome DevTools Protocol
                v
     Chrome/Chromium, dedicated profile, signed in once
                |
                v
        chatgpt.com (or any configured chat web UI)

The client never knows a browser is involved: it sends OpenAI requests and receives OpenAI responses, including SSE streaming.

What it does

  • POST /v1/chat/completions and POST /v1/completions (legacy), with stream: true.
  • GET /v1/models, GET /health, GET /v1/providers/{provider}/status, POST /v1/providers/{provider}/validate, POST /v1/admin/reload. Both provider routes accept the provider@account spelling, so one account of a provider can be checked on its own.
  • Multi-turn continuity: each OpenAI conversation is mapped to a web conversation, so only the newest message is sent once the thread exists.
  • Markdown fidelity: once the answer has settled the gateway clicks the provider's own copy button and reads the clipboard over CDP, so tables, code fences and lists arrive intact instead of being flattened by innerText.
  • Several providers: ChatGPT, Claude and DeepSeek ship as configuration files; adding another one is a TOML file, not a code change.
  • Several accounts per provider: each declared account gets its own browser profile, session and tab pool, and is selected with model = "chatgpt@perso".
  • Several completions per request: n > 1 is honoured, and the variants run in parallel when the account has more than one tab (max_tabs).
  • Optional local dashboard on /dashboard with live logs and manual tests.
  • Every selector lives in ~/.llm-gateway/providers/<provider>.toml, reloaded live: adapting to a UI change is a configuration edit, not a rebuild.

Requirements

  • Rust 1.85 or newer (tested with 1.98).
  • Chrome, Chromium, Edge or Brave. Firefox is not supported: it does not speak CDP.

Install and first run

cargo build --release
./target/release/llm-gateway config init      # creates ~/.llm-gateway
./target/release/llm-gateway login chatgpt    # sign in once, in the window that opens
./target/release/llm-gateway test chatgpt     # send one prompt and print the answer
./target/release/llm-gateway serve            # starts the API on 127.0.0.1:8080

Three providers are installed on the first run: chatgpt, claude and deepseek. Each needs its own sign-in:

./target/release/llm-gateway login claude
./target/release/llm-gateway selftest claude   # do the selectors still match?

The Claude and DeepSeek selectors were collected from public sources and have not been validated against a signed-in session here: run the selftest first, and see providers/selectors.md for which key to fix if the UI moved.

Then, with any OpenAI client:

# bash / zsh
curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model": "chatgpt", "messages": [{"role": "user", "content": "Bonjour"}]}'
# PowerShell: single quotes are literal, so do NOT escape the double quotes.
# Writing '{\"model\":...}' sends the backslashes and the body is rejected with
# "Failed to parse the request body as JSON".
curl.exe -s http://127.0.0.1:8080/v1/chat/completions -H "content-type: application/json" `
  -d '{"model":"chatgpt","messages":[{"role":"user","content":"Bonjour"}]}'

# Or, without any quoting puzzle at all:
$body = @{ model = "chatgpt"; messages = @(@{ role = "user"; content = "Bonjour" }) } |
        ConvertTo-Json -Depth 5
Invoke-RestMethod -Uri http://127.0.0.1:8080/v1/chat/completions -Method Post `
  -Body $body -ContentType application/json

# ./demo.ps1 walks through the whole API (health, models, streaming, status,
# validation, reload) so you do not have to type any of this by hand.
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="unused")
print(client.chat.completions.create(
    model="chatgpt",
    messages=[{"role": "user", "content": "Bonjour"}],
).choices[0].message.content)

Streaming works the same way with stream=True.

Command line

llm-gateway [GLOBAL OPTIONS] [COMMAND]

serve [--backend browser|mock]   start the API server (default command)
test <provider[@account]>|--all [--prompt TEXT] [--json]
selftest <provider[@account]>|--all [--json]
list [--json]
login <provider[@account]> [--account ID]   open a window to sign in once
logout <provider[@account]> [--account ID] [--all]   delete the stored profile(s)
conversations list|clear [--provider NAME[@ACCOUNT]] [--account ID]
reload                           ask a running server to reload its configuration
config path|init|show

Global options: --host, --port, --browser <chrome|chromium|edge|brave>, --browser-path, --profile-dir, --provider, --headless, --log-level, --log-format pretty|json, --config, --debug, --max-tabs, --api-key, --markdown auto|clipboard|dom|text, --yes.

Exit codes: 0 success, 1 runtime error, 2 selftest failure.

Model names

model selects the provider:

Value Meaning
chatgpt the provider, its default model
chatgpt/gpt-4o the provider and a model name recorded in the response
gpt-4o accepted when exactly one provider is configured

The model is echoed back but the web UI keeps whatever model it has selected; a warning is added to x_gateway_warnings when you ask for another one. Driving the UI model picker is not implemented yet.

When a provider declares several accounts, the account is part of the same string, between the provider and the model name:

Value Meaning
chatgpt the provider, its default account
chatgpt@perso the account whose id (or label) is perso
[email protected]@gmail.com the account with that address
chatgpt@perso/gpt-4o that account, and a model name

GET /v1/models lists one entry per provider and per declared account, so a client that only reads that endpoint still discovers them.

The same spelling works outside model, wherever a provider is named:

Command Effect
test chatgpt@perso / selftest chatgpt@perso checks that one account
test --all / selftest --all checks every account of every provider
GET /v1/providers/chatgpt@perso/status the state of that account
POST /v1/providers/chatgpt@pro/validate a full report for that account

A bare provider name means its default account, and the answer says which one was actually probed ("provider": "chatgpt@perso"), so two accounts of the same provider are never confused in a report or in the logs.

Multi-turn conversations

A client that resends the whole history (every OpenAI SDK does) is mapped to one web conversation:

  1. The conversation id comes from the X-Conversation-Id request header when present, otherwise from a fingerprint of the system prompt and the first user message (conversation.id_source = "fingerprint" in the config).
  2. The first turn replays the whole history into the page.
  3. Later turns reopen the stored web conversation and send only the newest message.
  4. The response always carries X-Conversation-Id, so a client can pin a thread explicitly.
  5. If the web conversation is gone, the turn automatically falls back to replaying the history in a new one.

~/.llm-gateway/conversations.json is plain JSON you can read and edit:

{
  "version": 1,
  "threads": {
    "fp-3f1a...": {
      "provider": "chatgpt",
      "account": "perso",
      "web_url": "https://chatgpt.com/c/68f0...",
      "web_id": "68f0...",
      "created_at": 1760000000,
      "last_used_at": 1760000100,
      "turns": 3,
      "messages_sent": 7,
      "state": "ready"
    }
  }
}

llm-gateway conversations list and clear inspect and prune it. The listing shows the account that served each thread, and both commands accept --provider chatgpt@perso or --account perso to narrow the selection; --account default selects the threads of a provider that declares no account. A thread is marked incomplete when a client disconnects mid-generation; the next turn then replays the history instead of trusting a half-written thread.

OpenAI parameter support

Parameter Behaviour
model selects the provider (see above)
messages string content or an array of text / image_url parts
stream SSE chunks, data: [DONE] termination
stream_options.include_usage adds a final usage chunk
temperature, top_p, max_tokens, stop, response_format, seed, user, penalties, logprobs ignored, logged, and listed in x_gateway_warnings
tools, functions, tool_choice not supported: HTTP 400 unsupported_parameter
n > 1 supported: n completions, each in its own web conversation. Clamped to capture.max_variants (4 by default) with a warning. Variants run in parallel when the account allows more than one tab

usage is estimated with tiktoken when its vocabulary can be loaded, and with a characters/4 heuristic otherwise (offline machines).

/v1/completions accepts a string or an array prompt and answers with the legacy text_completion shape. Streaming is not available on that endpoint.

Signing in with Google (or any OAuth provider that blocks automation)

Symptom: after llm-gateway login chatgpt, the Google window answers "This browser or app may not be secure".

Why it happens: Google refuses to sign a browser it identifies as automated or outdated. Two things mattered here, and the first one was a bug in this project:

  1. The stealth helper used to advertise an outdated Chrome user agent (Chrome/107 in 2026). That alone reads as "old browser = insecure", which is exactly what the message means. It is fixed: the default is now [browser] stealth = "off" and the user agent is never spoofed.
  2. Google can also refuse a browser it knows is being driven over CDP. That part cannot be argued with, so the gateway offers a sign-in path where no automation is involved at all.
llm-gateway login chatgpt --manual      # prints the exact command, or add --launch

It tells you to run something like:

"C:\Program Files\Google\Chrome\Application\chrome.exe" \
  --user-data-dir="C:\Users\you\.llm-gateway\profiles\chatgpt\user-data-dir" \
  --no-first-run --no-default-browser-check https://chatgpt.com/

Sign in with Google in that window: it is an ordinary Chrome launched by you, with no automation switch and no debugging port. Then close it and press Enter. The gateway relaunches that same profile for later requests, so the session is already there.

Alternative: keep that window and let the gateway drive it

Start the same command with --remote-debugging-port=9222 (the --manual output prints it) and keep it open, then:

llm-gateway serve --attach --debug-port 9222

The gateway attaches to your browser, reuses a tab it can see (otherwise it opens one in that same browser, with your session), and never closes your browser when it shuts down. If Google still refuses the sign-in with the debugging port open, use the recommended path above.

Also worth knowing

  • Email + password sign-in on the provider site is not blocked by any of this; only Google's OAuth screen applies these checks.
  • --stealth minimal adds a navigator.webdriver patch on top of the default; --stealth aggressive restores the old user-agent spoofing and is documented as harmful for sign-in.
  • When a request hits a signed-out page, the API answers with 401 requires_login and the message now mentions the --manual path.

Configuration

~/.llm-gateway/config.toml (created on first run, see config.toml.example for the annotated version). Precedence: CLI flags > LLM_GATEWAY_* environment variables > the file > defaults.

The settings that matter most:

[browser]
executable = ""      # empty: auto-detect Chrome, Edge or Brave
headless = false     # keep it false so you can sign in and solve captchas
max_tabs = 1         # concurrent turns per provider *account*, i.e. tabs kept open
busy_wait_s = 30     # how long a request waits for a free tab before a 429

[conversation]
strategy = "auto"    # auto | reuse | replay
id_source = "fingerprint"  # fingerprint | header | uuid

[capture]
poll_interval_ms = 400   # DOM polling while waiting for the answer
quiet_ms = 1500          # silence after which the answer is considered final
response_timeout_s = 180 # hard budget for one answer
markdown = "auto"        # auto | clipboard | dom | text (see below)
max_variants = 4         # upper bound accepted for the "n" parameter

max_tabs is the setting that turns the gateway from a queue into a pool: with max_tabs = 2 an account can answer two requests (or two variants of one n = 2 request) at the same time, in two tabs of the same browser.

capture.markdown decides where the answer text comes from:

Value Behaviour
auto the provider's copy button (Markdown through the clipboard) when it has one, the DOM conversion otherwise
clipboard always the copy button; falls back to the DOM when the clipboard cannot be read
dom always convert the answer HTML to Markdown
text the visible text only, formatting flattened (the historical behaviour)

A provider may override it in its own [input] section. Streamed deltas always stay plain text: only the settled answer is re-read with full fidelity.

Provider files (~/.llm-gateway/providers/*.toml) hold the URLs, the CSS selectors, the timeouts and the capabilities of each web UI. They are watched: editing one reloads it within a second, without restarting the server or the browser. See providers/selectors.md.

They are copies created on the first run, so upgrading the binary does not change them. Refresh them after an upgrade with llm-gateway config init --force (the previous file is kept next to it as <provider>.toml.bak).

llm-gateway --version prints the compilation date (0.1.0 (built 2026-09-18T17:26:11Z)): if it does not match your sources, run cargo build --release again.

Error mapping

Errors use the OpenAI error shape and stable codes:

Situation HTTP code
Unknown model 404 model_not_found
Unknown account in the model string 404 account_not_found (the message lists the known ones)
Bad request, ignored-but-unsupported parameter 400 invalid_request, unsupported_parameter
Images sent to a provider without upload support 400 unsupported_content_type
Missing or wrong API key 401 invalid_api_key
Signed out 401 requires_login
Captcha challenge 403 captcha_required
Browser profile locked by another instance 409 browser_profile_locked
Tab busy (30 s wait) or upstream throttling 429 provider_busy, upstream_rate_limit
Selector gone, upstream error 500 selector_missing, upstream_error
Browser missing, provider misconfigured 503 browser_unavailable, provider_misconfigured
No answer in time 504 upstream_timeout

On a failure the server saves a screenshot and the full DOM in ~/.llm-gateway/debug/, plus one line per failure in debug/index.jsonl (--debug does it for every request).

Dashboard

http://127.0.0.1:8080/dashboard shows the providers, their live state, the recent logs (SSE) and lets you send a prompt by hand. If an API key is configured, paste it in the header field.

Development

cargo fmt --all --check
cargo clippy --all-targets -- -D warnings
cargo test                                    # no browser, no network
./scripts/e2e.ps1                             # real Chrome, local fixture page

./scripts/ci.ps1                              # all of the above, in one go
./scripts/ci.ps1 -E2e                         # ... plus the browser suite

.github/workflows/ci.yml runs the same gates on Linux and Windows for every push and pull request, and builds the release binary. The browser suite needs a real Chromium, so it is a separate job that runs on demand (or nightly) instead of blocking a machine that has no browser.

  • tests/api_integration.rs drives the whole HTTP surface through a mock backend, including multi-account routing and n > 1.
  • tests/e2e_browser.rs drives a real Chrome against tests/fixtures/chatgpt_mock.html, a static fake of the ChatGPT DOM with login, captcha, rate-limit, error, timeout, rewrite and attachment modes. That fixture is what makes the capture loop regression-testable without a network or an account.
  • demo.ps1 / demo.sh run a full cycle; add -Mock (or --mock) to run without a browser or a login.

Limits and honest caveats

  • Streamed deltas are best effort. The final answer written into a non-streaming response is always the exact captured text. If the web UI rewrites text it has already shown (React re-render, code block highlighting), a streaming client may have received text that the final answer no longer contains. A warning is logged when that happens.
  • Markdown fidelity depends on the provider. The clipboard path is exact, because the provider copies what it rendered; the DOM conversion is a faithful approximation that covers the constructs a chat UI actually produces. A page that refuses the clipboard read (or a provider with no copy button) silently falls back to it. Only the settled answer is re-read: streamed deltas stay plain text.
  • One turn at a time per provider account by default: a second request waits up to browser.busy_wait_s (30 s) and then gets a 429. Raise max_tabs to let an account answer several requests at once, in several tabs.
  • Only one tab per browser is rendered by Chrome. On the other tabs the prompt is submitted with in-page events rather than a synthetic mouse click, which every chat UI tested here accepts, and the clipboard read still works because the permission is granted over CDP.
  • Selectors are the fragile part. The fixture page covers regressions in our capture loop; it cannot predict a change on the real site. Run llm-gateway selftest after a UI update. The Claude and DeepSeek files are starting points collected from public sources, not validated against a signed-in session.
  • No tool calling, no model switching.
  • Latency is browser latency: a fresh conversation takes a few seconds to start generating, plus the page load.

Security and terms of use

  • The API binds to 127.0.0.1 by default. If you expose it on a network, configure server.api_key (or --api-key) and put it behind TLS.
  • Browser profiles contain live session cookies: they live in ~/.llm-gateway/profiles/<provider> and are never encrypted by this tool. Use disk encryption and llm-gateway logout --all when you are done.
  • The tool only uses your own signed-in subscriptions. It does not bypass paywalls, quotas or captchas: when a challenge appears it stops and tells you.
  • You remain responsible for complying with each provider's terms of service. Automating a web UI may be restricted by them; check before you rely on it.

Layout of the state directory

~/.llm-gateway/
  config.toml                  global configuration
  providers/                   one TOML per provider (selectors, URLs, timeouts)
  profiles/<provider>/user-data-dir/            persistent Chrome profile
  profiles/<provider>/accounts/<account>/user-data-dir/   one per declared account
  conversations.json           client conversation -> web conversation mapping
  debug/                       screenshots, DOM dumps, index.jsonl
  logs/llm-gateway.log         JSON logs

Providers and accounts

Provider File Models Notes
OpenAI ChatGPT providers/chatgpt.toml gpt-4o, gpt-4o-mini, gpt-4.1 #prompt-textarea, per-answer copy button
Anthropic Claude providers/claude.toml claude-sonnet-4, claude-opus-4, claude-haiku-4 ProseMirror editor; selectors from public sources
DeepSeek providers/deepseek.toml deepseek-chat, deepseek-reasoner textarea#chat-input, .ds-markdown answers

An account is a browser profile: declaring two ChatGPT accounts gives two independent sessions, each with its own tab pool, driven side by side.

[[accounts]]
id = "perso"
email = "[email protected]"
default = true

[[accounts]]
id = "pro"
email = "[email protected]"
llm-gateway login chatgpt --account perso
llm-gateway login chatgpt@pro
llm-gateway list                 # shows every account and its profile
llm-gateway logout chatgpt@pro   # forget one account

The model string then carries the account: {"model": "chatgpt@pro"}, or {"model": "[email protected]@gmail.com"}, or {"model": "chatgpt@pro/gpt-4o"}. Each account keeps its own thread store entries, so the same prompt sent to two accounts never shares one web conversation. Adding accounts to a provider that was already signed in changes its profile layout (profiles/<provider>/accounts/<id>/…), so sign that provider in again once.