llm-bridge — llm-gateway
A local, OpenAI-compatible HTTP API that answers by driving a real, signed-in browser session on an LLM web UI instead of a paid API key.
OpenAI SDK / curl / Continue / Cursor
|
| POST /v1/chat/completions (OpenAI format)
v
llm-gateway (Axum, Rust)
|
| Chrome DevTools Protocol
v
Chrome/Chromium, dedicated profile, signed in once
|
v
chatgpt.com (or any configured chat web UI)
The client never knows a browser is involved: it sends OpenAI requests and receives OpenAI responses, including SSE streaming.
What it does
POST /v1/chat/completionsandPOST /v1/completions(legacy), withstream: true.GET /v1/models,GET /health,GET /v1/providers/{provider}/status,POST /v1/providers/{provider}/validate,POST /v1/admin/reload. Both provider routes accept theprovider@accountspelling, so one account of a provider can be checked on its own.- Multi-turn continuity: each OpenAI conversation is mapped to a web conversation, so only the newest message is sent once the thread exists.
- Markdown fidelity: once the answer has settled the gateway clicks the
provider's own copy button and reads the clipboard over CDP, so tables, code
fences and lists arrive intact instead of being flattened by
innerText. - Several providers: ChatGPT, Claude and DeepSeek ship as configuration files; adding another one is a TOML file, not a code change.
- Several accounts per provider: each declared account gets its own browser
profile, session and tab pool, and is selected with
model = "chatgpt@perso". - Several completions per request:
n > 1is honoured, and the variants run in parallel when the account has more than one tab (max_tabs). - Optional local dashboard on
/dashboardwith live logs and manual tests. - Every selector lives in
~/.llm-gateway/providers/<provider>.toml, reloaded live: adapting to a UI change is a configuration edit, not a rebuild.
Requirements
- Rust 1.85 or newer (tested with 1.98).
- Chrome, Chromium, Edge or Brave. Firefox is not supported: it does not speak CDP.
Install and first run
cargo build --release
./target/release/llm-gateway config init # creates ~/.llm-gateway
./target/release/llm-gateway login chatgpt # sign in once, in the window that opens
./target/release/llm-gateway test chatgpt # send one prompt and print the answer
./target/release/llm-gateway serve # starts the API on 127.0.0.1:8080
Three providers are installed on the first run: chatgpt, claude and
deepseek. Each needs its own sign-in:
./target/release/llm-gateway login claude
./target/release/llm-gateway selftest claude # do the selectors still match?
The Claude and DeepSeek selectors were collected from public sources and have not
been validated against a signed-in session here: run the selftest first, and see
providers/selectors.md for which key to fix if the UI moved.
Then, with any OpenAI client:
# bash / zsh
curl -s http://127.0.0.1:8080/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model": "chatgpt", "messages": [{"role": "user", "content": "Bonjour"}]}'
# PowerShell: single quotes are literal, so do NOT escape the double quotes.
# Writing '{\"model\":...}' sends the backslashes and the body is rejected with
# "Failed to parse the request body as JSON".
curl.exe -s http://127.0.0.1:8080/v1/chat/completions -H "content-type: application/json" `
-d '{"model":"chatgpt","messages":[{"role":"user","content":"Bonjour"}]}'
# Or, without any quoting puzzle at all:
$body = @{ model = "chatgpt"; messages = @(@{ role = "user"; content = "Bonjour" }) } |
ConvertTo-Json -Depth 5
Invoke-RestMethod -Uri http://127.0.0.1:8080/v1/chat/completions -Method Post `
-Body $body -ContentType application/json
# ./demo.ps1 walks through the whole API (health, models, streaming, status,
# validation, reload) so you do not have to type any of this by hand.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="unused")
print(client.chat.completions.create(
model="chatgpt",
messages=[{"role": "user", "content": "Bonjour"}],
).choices[0].message.content)
Streaming works the same way with stream=True.
Command line
llm-gateway [GLOBAL OPTIONS] [COMMAND]
serve [--backend browser|mock] start the API server (default command)
test <provider[@account]>|--all [--prompt TEXT] [--json]
selftest <provider[@account]>|--all [--json]
list [--json]
login <provider[@account]> [--account ID] open a window to sign in once
logout <provider[@account]> [--account ID] [--all] delete the stored profile(s)
conversations list|clear [--provider NAME[@ACCOUNT]] [--account ID]
reload ask a running server to reload its configuration
config path|init|show
Global options: --host, --port, --browser <chrome|chromium|edge|brave>,
--browser-path, --profile-dir, --provider, --headless, --log-level,
--log-format pretty|json, --config, --debug, --max-tabs, --api-key,
--markdown auto|clipboard|dom|text, --yes.
Exit codes: 0 success, 1 runtime error, 2 selftest failure.
Model names
model selects the provider:
| Value | Meaning |
|---|---|
chatgpt |
the provider, its default model |
chatgpt/gpt-4o |
the provider and a model name recorded in the response |
gpt-4o |
accepted when exactly one provider is configured |
The model is echoed back but the web UI keeps whatever model it has
selected; a warning is added to x_gateway_warnings when you ask for another
one. Driving the UI model picker is not implemented yet.
When a provider declares several accounts, the account is part of the same string, between the provider and the model name:
| Value | Meaning |
|---|---|
chatgpt |
the provider, its default account |
chatgpt@perso |
the account whose id (or label) is perso |
[email protected]@gmail.com |
the account with that address |
chatgpt@perso/gpt-4o |
that account, and a model name |
GET /v1/models lists one entry per provider and per declared account, so a
client that only reads that endpoint still discovers them.
The same spelling works outside model, wherever a provider is named:
| Command | Effect |
|---|---|
test chatgpt@perso / selftest chatgpt@perso |
checks that one account |
test --all / selftest --all |
checks every account of every provider |
GET /v1/providers/chatgpt@perso/status |
the state of that account |
POST /v1/providers/chatgpt@pro/validate |
a full report for that account |
A bare provider name means its default account, and the answer says which one was
actually probed ("provider": "chatgpt@perso"), so two accounts of the same
provider are never confused in a report or in the logs.
Multi-turn conversations
A client that resends the whole history (every OpenAI SDK does) is mapped to one web conversation:
- The conversation id comes from the
X-Conversation-Idrequest header when present, otherwise from a fingerprint of the system prompt and the first user message (conversation.id_source = "fingerprint"in the config). - The first turn replays the whole history into the page.
- Later turns reopen the stored web conversation and send only the newest message.
- The response always carries
X-Conversation-Id, so a client can pin a thread explicitly. - If the web conversation is gone, the turn automatically falls back to replaying the history in a new one.
~/.llm-gateway/conversations.json is plain JSON you can read and edit:
{
"version": 1,
"threads": {
"fp-3f1a...": {
"provider": "chatgpt",
"account": "perso",
"web_url": "https://chatgpt.com/c/68f0...",
"web_id": "68f0...",
"created_at": 1760000000,
"last_used_at": 1760000100,
"turns": 3,
"messages_sent": 7,
"state": "ready"
}
}
}
llm-gateway conversations list and clear inspect and prune it. The listing
shows the account that served each thread, and both commands accept
--provider chatgpt@perso or --account perso to narrow the selection;
--account default selects the threads of a provider that declares no account.
A thread is marked incomplete when a client disconnects mid-generation; the
next turn then replays the history instead of trusting a half-written thread.
OpenAI parameter support
| Parameter | Behaviour |
|---|---|
model |
selects the provider (see above) |
messages |
string content or an array of text / image_url parts |
stream |
SSE chunks, data: [DONE] termination |
stream_options.include_usage |
adds a final usage chunk |
temperature, top_p, max_tokens, stop, response_format, seed, user, penalties, logprobs |
ignored, logged, and listed in x_gateway_warnings |
tools, functions, tool_choice |
not supported: HTTP 400 unsupported_parameter |
n > 1 |
supported: n completions, each in its own web conversation. Clamped to capture.max_variants (4 by default) with a warning. Variants run in parallel when the account allows more than one tab |
usage is estimated with tiktoken when its vocabulary can be loaded, and with
a characters/4 heuristic otherwise (offline machines).
/v1/completions accepts a string or an array prompt and answers with the
legacy text_completion shape. Streaming is not available on that endpoint.
Signing in with Google (or any OAuth provider that blocks automation)
Symptom: after llm-gateway login chatgpt, the Google window answers
"This browser or app may not be secure".
Why it happens: Google refuses to sign a browser it identifies as automated or outdated. Two things mattered here, and the first one was a bug in this project:
- The stealth helper used to advertise an outdated Chrome user agent
(
Chrome/107in 2026). That alone reads as "old browser = insecure", which is exactly what the message means. It is fixed: the default is now[browser] stealth = "off"and the user agent is never spoofed. - Google can also refuse a browser it knows is being driven over CDP. That part cannot be argued with, so the gateway offers a sign-in path where no automation is involved at all.
Recommended: sign in from a browser you start yourself
llm-gateway login chatgpt --manual # prints the exact command, or add --launch
It tells you to run something like:
"C:\Program Files\Google\Chrome\Application\chrome.exe" \
--user-data-dir="C:\Users\you\.llm-gateway\profiles\chatgpt\user-data-dir" \
--no-first-run --no-default-browser-check https://chatgpt.com/
Sign in with Google in that window: it is an ordinary Chrome launched by you, with no automation switch and no debugging port. Then close it and press Enter. The gateway relaunches that same profile for later requests, so the session is already there.
Alternative: keep that window and let the gateway drive it
Start the same command with --remote-debugging-port=9222 (the --manual output
prints it) and keep it open, then:
llm-gateway serve --attach --debug-port 9222
The gateway attaches to your browser, reuses a tab it can see (otherwise it opens one in that same browser, with your session), and never closes your browser when it shuts down. If Google still refuses the sign-in with the debugging port open, use the recommended path above.
Also worth knowing
- Email + password sign-in on the provider site is not blocked by any of this; only Google's OAuth screen applies these checks.
--stealth minimaladds anavigator.webdriverpatch on top of the default;--stealth aggressiverestores the old user-agent spoofing and is documented as harmful for sign-in.- When a request hits a signed-out page, the API answers with
401 requires_loginand the message now mentions the--manualpath.
Configuration
~/.llm-gateway/config.toml (created on first run, see
config.toml.example for the annotated version). Precedence:
CLI flags > LLM_GATEWAY_* environment variables > the file > defaults.
The settings that matter most:
[browser]
executable = "" # empty: auto-detect Chrome, Edge or Brave
headless = false # keep it false so you can sign in and solve captchas
max_tabs = 1 # concurrent turns per provider *account*, i.e. tabs kept open
busy_wait_s = 30 # how long a request waits for a free tab before a 429
[conversation]
strategy = "auto" # auto | reuse | replay
id_source = "fingerprint" # fingerprint | header | uuid
[capture]
poll_interval_ms = 400 # DOM polling while waiting for the answer
quiet_ms = 1500 # silence after which the answer is considered final
response_timeout_s = 180 # hard budget for one answer
markdown = "auto" # auto | clipboard | dom | text (see below)
max_variants = 4 # upper bound accepted for the "n" parameter
max_tabs is the setting that turns the gateway from a queue into a pool: with
max_tabs = 2 an account can answer two requests (or two variants of one n = 2
request) at the same time, in two tabs of the same browser.
capture.markdown decides where the answer text comes from:
| Value | Behaviour |
|---|---|
auto |
the provider's copy button (Markdown through the clipboard) when it has one, the DOM conversion otherwise |
clipboard |
always the copy button; falls back to the DOM when the clipboard cannot be read |
dom |
always convert the answer HTML to Markdown |
text |
the visible text only, formatting flattened (the historical behaviour) |
A provider may override it in its own [input] section. Streamed deltas always
stay plain text: only the settled answer is re-read with full fidelity.
Provider files (~/.llm-gateway/providers/*.toml) hold the URLs, the CSS
selectors, the timeouts and the capabilities of each web UI. They are watched:
editing one reloads it within a second, without restarting the server or the
browser. See providers/selectors.md.
They are copies created on the first run, so upgrading the binary does not
change them. Refresh them after an upgrade with
llm-gateway config init --force (the previous file is kept next to it as
<provider>.toml.bak).
llm-gateway --version prints the compilation date
(0.1.0 (built 2026-09-18T17:26:11Z)): if it does not match your sources, run
cargo build --release again.
Error mapping
Errors use the OpenAI error shape and stable codes:
| Situation | HTTP | code |
|---|---|---|
| Unknown model | 404 | model_not_found |
| Unknown account in the model string | 404 | account_not_found (the message lists the known ones) |
| Bad request, ignored-but-unsupported parameter | 400 | invalid_request, unsupported_parameter |
| Images sent to a provider without upload support | 400 | unsupported_content_type |
| Missing or wrong API key | 401 | invalid_api_key |
| Signed out | 401 | requires_login |
| Captcha challenge | 403 | captcha_required |
| Browser profile locked by another instance | 409 | browser_profile_locked |
| Tab busy (30 s wait) or upstream throttling | 429 | provider_busy, upstream_rate_limit |
| Selector gone, upstream error | 500 | selector_missing, upstream_error |
| Browser missing, provider misconfigured | 503 | browser_unavailable, provider_misconfigured |
| No answer in time | 504 | upstream_timeout |
On a failure the server saves a screenshot and the full DOM in
~/.llm-gateway/debug/, plus one line per failure in debug/index.jsonl
(--debug does it for every request).
Dashboard
http://127.0.0.1:8080/dashboard shows the providers, their live state, the
recent logs (SSE) and lets you send a prompt by hand. If an API key is
configured, paste it in the header field.
Development
cargo fmt --all --check
cargo clippy --all-targets -- -D warnings
cargo test # no browser, no network
./scripts/e2e.ps1 # real Chrome, local fixture page
./scripts/ci.ps1 # all of the above, in one go
./scripts/ci.ps1 -E2e # ... plus the browser suite
.github/workflows/ci.yml runs the same gates on Linux and Windows for every
push and pull request, and builds the release binary. The browser suite needs a
real Chromium, so it is a separate job that runs on demand (or nightly) instead
of blocking a machine that has no browser.
tests/api_integration.rsdrives the whole HTTP surface through a mock backend, including multi-account routing andn > 1.tests/e2e_browser.rsdrives a real Chrome againsttests/fixtures/chatgpt_mock.html, a static fake of the ChatGPT DOM with login, captcha, rate-limit, error, timeout, rewrite and attachment modes. That fixture is what makes the capture loop regression-testable without a network or an account.demo.ps1/demo.shrun a full cycle; add-Mock(or--mock) to run without a browser or a login.
Limits and honest caveats
- Streamed deltas are best effort. The final answer written into a non-streaming response is always the exact captured text. If the web UI rewrites text it has already shown (React re-render, code block highlighting), a streaming client may have received text that the final answer no longer contains. A warning is logged when that happens.
- Markdown fidelity depends on the provider. The clipboard path is exact, because the provider copies what it rendered; the DOM conversion is a faithful approximation that covers the constructs a chat UI actually produces. A page that refuses the clipboard read (or a provider with no copy button) silently falls back to it. Only the settled answer is re-read: streamed deltas stay plain text.
- One turn at a time per provider account by default: a second request waits
up to
browser.busy_wait_s(30 s) and then gets a 429. Raisemax_tabsto let an account answer several requests at once, in several tabs. - Only one tab per browser is rendered by Chrome. On the other tabs the prompt is submitted with in-page events rather than a synthetic mouse click, which every chat UI tested here accepts, and the clipboard read still works because the permission is granted over CDP.
- Selectors are the fragile part. The fixture page covers regressions in our
capture loop; it cannot predict a change on the real site. Run
llm-gateway selftestafter a UI update. The Claude and DeepSeek files are starting points collected from public sources, not validated against a signed-in session. - No tool calling, no model switching.
- Latency is browser latency: a fresh conversation takes a few seconds to start generating, plus the page load.
Security and terms of use
- The API binds to
127.0.0.1by default. If you expose it on a network, configureserver.api_key(or--api-key) and put it behind TLS. - Browser profiles contain live session cookies: they live in
~/.llm-gateway/profiles/<provider>and are never encrypted by this tool. Use disk encryption andllm-gateway logout --allwhen you are done. - The tool only uses your own signed-in subscriptions. It does not bypass paywalls, quotas or captchas: when a challenge appears it stops and tells you.
- You remain responsible for complying with each provider's terms of service. Automating a web UI may be restricted by them; check before you rely on it.
Layout of the state directory
~/.llm-gateway/
config.toml global configuration
providers/ one TOML per provider (selectors, URLs, timeouts)
profiles/<provider>/user-data-dir/ persistent Chrome profile
profiles/<provider>/accounts/<account>/user-data-dir/ one per declared account
conversations.json client conversation -> web conversation mapping
debug/ screenshots, DOM dumps, index.jsonl
logs/llm-gateway.log JSON logs
Providers and accounts
| Provider | File | Models | Notes |
|---|---|---|---|
| OpenAI ChatGPT | providers/chatgpt.toml |
gpt-4o, gpt-4o-mini, gpt-4.1 |
#prompt-textarea, per-answer copy button |
| Anthropic Claude | providers/claude.toml |
claude-sonnet-4, claude-opus-4, claude-haiku-4 |
ProseMirror editor; selectors from public sources |
| DeepSeek | providers/deepseek.toml |
deepseek-chat, deepseek-reasoner |
textarea#chat-input, .ds-markdown answers |
An account is a browser profile: declaring two ChatGPT accounts gives two independent sessions, each with its own tab pool, driven side by side.
[[accounts]]
id = "perso"
email = "[email protected]"
default = true
[[accounts]]
id = "pro"
email = "[email protected]"
llm-gateway login chatgpt --account perso
llm-gateway login chatgpt@pro
llm-gateway list # shows every account and its profile
llm-gateway logout chatgpt@pro # forget one account
The model string then carries the account:
{"model": "chatgpt@pro"}, or {"model": "[email protected]@gmail.com"},
or {"model": "chatgpt@pro/gpt-4o"}. Each account keeps its own thread store
entries, so the same prompt sent to two accounts never shares one web
conversation. Adding accounts to a provider that was already signed in changes
its profile layout (profiles/<provider>/accounts/<id>/…), so sign that
provider in again once.