Files
bruno 7b5aa6fcc9
CI / fmt, clippy and tests (ubuntu-latest) (push) Successful in 21m0s
CI / browser end-to-end (Chromium + fixture) (push) Skipped
CI / fmt, clippy and tests (windows-latest) (push) Canceled after 0s
Add initial llm-bridge project
2026-09-18 18:01:07 -04:00

170 lines
9.3 KiB
Markdown

# Provider selectors
Every selector lives in a provider configuration file
(`~/.llm-gateway/providers/<provider>.toml`), never in the Rust source. When a
web UI changes, edit the TOML: the server picks the change up within a second
(no restart, no browser relaunch). `llm-gateway selftest <provider>` tells you
whether the current selectors still work.
## Why this file exists
Web UIs change without notice. The three failure modes we care about are:
| Symptom | Reported as | What to do |
|---|---|---|
| The prompt field disappeared | HTTP 500, code `selector_missing` | Update `selectors.input_field` |
| The send button disappeared | HTTP 500, code `selector_missing` | Update `selectors.send_button` |
| The answer container moved | empty answer or HTTP 500 | Update `selectors.response_container` |
| The streaming indicator is wrong | answers truncated or 504 | Update `selectors.streaming_indicator` |
| The session expired | HTTP 401, code `requires_login` | Run `llm-gateway login <provider>` |
A failing capture always leaves a screenshot and a DOM dump in
`~/.llm-gateway/debug/`, plus one line per failure in `debug/index.jsonl`.
## How to update a selector
1. Open the provider page in a normal browser, signed in.
2. Open the developer tools, inspect the element you need.
3. Prefer stable hooks in this order: `data-testid` > `id` > semantic attributes
(`aria-label`, `role`, `type`) > a CSS class. Avoid generated class names such
as `.css-1x2y3z`: they change on every deploy.
4. Put a comma separated list in the TOML when a UI has several variants
(`div#prompt-textarea, div.ProseMirror[contenteditable='true']`): the first
match wins and the fallbacks keep you running through a partial rollout.
5. Run `llm-gateway selftest <provider>` and keep the TOML only when it reports
`ok`.
## Required and optional selectors
| Key | Required | Used for |
|---|---|---|
| `input_field` | yes | focusing the editor, verifying the injected text, detecting an empty editor after submit |
| `send_button` | yes | submitting the prompt (Enter is the fallback) |
| `response_container` | yes | one node per assistant answer, in document order; the driver reads the n-th node. It must select the **answer text**, not the whole exchange: a container that also holds the action toolbar would capture the word "Copy" as part of the answer |
| `copy_button` | no | the per-answer "copy" control, which gives the answer as Markdown through the clipboard. It may be inside the answer node or a sibling of it: the driver looks for it there and never elsewhere on the page, so the copy button of a previous answer is never clicked |
| `streaming_indicator` | recommended | element that only exists while the answer is generated (a stop button) |
| `file_input` | for images | hidden input fed through `DOM.setFileInputFiles` |
| `attachment_indicator` | optional | chip shown once an upload has been ingested |
| `new_chat_button` | optional | starting a brand new conversation without reloading |
| `login_page_indicator` | recommended | detecting a sign-in page even when the URL does not say so. On chatgpt.com the signed-out landing page exposes `button[data-testid='login']` |
| `[login] text_patterns` | recommended | lowercase substrings of the visible text that reveal a signed-out page; only consulted when the prompt field is missing, so a chat page that mentions "log in" is never misclassified |
| `captcha_indicator` | recommended | detecting a bot challenge |
| `error_banner` | recommended | surfacing the upstream error message to the API client |
| `model_selector` | reserved | not driven yet: the gateway never changes the model in the UI |
## Answer fidelity (Markdown)
A web UI renders its answer as rich HTML and `innerText` flattens it: a table
collapses into loose lines and a code block loses its fences. Once the answer has
settled the driver therefore reads it a second time, in this order:
1. **the clipboard** — it clicks the provider's own copy button and reads the
clipboard over CDP (`Browser.setPermission` then
`navigator.clipboard.readText()`). The provider copies exactly the Markdown it
rendered, so nothing is guessed. Requires `selectors.copy_button`.
2. **the DOM** — it converts the answer node's `innerHTML` to Markdown
(headings, lists, quotes, links, images, fenced code, GitHub tables).
3. **the visible text** — the historical behaviour, used when both fail.
`[capture] markdown` chooses between them and a provider may override it in its
own `[input]` section. Streamed deltas always come from the DOM, so they stay
plain text; the final answer — the one a non-streaming client receives — is the
converted one, and the source used is logged as `fidelity`.
## Current status
| Provider | File | Last validated | Status |
|---|---|---|---|
| OpenAI ChatGPT | `providers/chatgpt.toml` | 2026-09-18, signed-in session | `selftest` ok; `fidelity = "clipboard"` |
| Anthropic Claude | `providers/claude.toml` | 2026-09-18, signed-in session | `selftest` ok; `fidelity = "clipboard"` |
| DeepSeek | `providers/deepseek.toml` | 2026-09-18, signed-in session | `selftest` ok; **`fidelity = "dom"`**: the copy button has not been observed yet |
| Local fixture (tests only) | generated by `tests/e2e_browser.rs` | 2026-09-18 | validated: full pipeline, streaming, login, captcha, rate limit, error banner, timeout, missing selector, rewrite, attachment, Markdown fidelity |
The Claude and DeepSeek selectors were collected from public sources rather than
observed on a signed-in session. Treat them as a starting point: the first run of
`llm-gateway selftest <provider>` tells you whether they still match, and the
failure modes above tell you which key to fix.
The ChatGPT selectors shipped in this repository are the ones observed on the
public web UI in September 2026. They are a starting point, not a guarantee:
run the selftest before relying on them.
## Upgrading the installed files
The provider files under `~/.llm-gateway/providers/` are **copies** made on the
first run: editing the ones in this repository does not change them. After
upgrading the binary, refresh them with:
```bash
llm-gateway config init --force # keeps the previous file as <provider>.toml.bak
```
## Localised user interfaces
A provider translates its own UI, and the browser locale of the machine does not
change that: the same ChatGPT account serves a French page under `lang="fr-CA"`
and an English one under `lang="en-US"`. Every `aria-label` fallback therefore
only matches one language, which is why the stable hooks come first:
| Observed (French UI, September 2026) | Selector to prefer |
|---|---|
| `<button type="submit" id="composer-submit-button" aria-label="Envoyer la requête" data-testid="send-button">` | `button[data-testid='send-button']`, `button#composer-submit-button` |
| the answer toolbar copy control | `button[data-testid='copy-turn-action-button']` |
| `<textarea placeholder="Message DeepSeek" name="search" rows="2">` | `textarea[placeholder*='Message DeepSeek']` - the English placeholder and the `#chat-input` id of older builds are gone |
The `[login] text_patterns` follow the same rule: they are lowercase substrings of
the page text, so a French sign-in page needs French patterns
(`"se connecter"`, `"s'inscrire"`) next to the English ones. A page that is not
recognised as a sign-in page is reported as a missing prompt field instead, which
is a confusing way to learn that the session expired.
## Several accounts of one provider
A signed-in account is a browser profile, so driving two ChatGPT accounts means
two profiles, two sessions and two tab pools. Declare them in the provider file:
```toml
[[accounts]]
id = "perso" # the short name used in the model string
email = "[email protected]" # also accepted as the account key
label = "Personal" # free-form, shown by 'llm-gateway list'
default = true # used by {"model": "chatgpt"}
[[accounts]]
id = "pro"
email = "[email protected]"
```
Then sign each one in once and address it in the model string:
```bash
llm-gateway login chatgpt --account perso
llm-gateway login chatgpt@pro # the same thing
```
```jsonc
{"model": "chatgpt"} // the default account
{"model": "chatgpt@pro"} // by id
{"model": "[email protected]@gmail.com"} // by address
{"model": "chatgpt@pro/gpt-4o"} // account and model
```
Their profiles live in `profiles/<provider>/accounts/<account>/user-data-dir`.
A provider that declares no account keeps the historical
`profiles/<provider>/user-data-dir`, so adding accounts later means signing in
again for the first one. Each account has its own tab pool, so `max_tabs`
applies per account.
## Adding a new provider
1. Copy `providers/chatgpt.toml` to `~/.llm-gateway/providers/<name>.toml`.
2. Set `provider.name`, `provider.web_url`, `provider.new_conversation_url` and
`provider.conversation_url_pattern` (one capture group, the web conversation
id extracted from the page URL).
3. Fill in the selectors, then `llm-gateway selftest <name>`.
4. `llm-gateway login <name>` once, and the provider is usable as a model:
`{"model": "<name>"}` or `{"model": "<name>/<model>"}`.
Nothing else is provider specific: the capture loop, the error mapping, the
conversation store and the API are shared.