170 lines
9.3 KiB
Markdown
170 lines
9.3 KiB
Markdown
# Provider selectors
|
|
|
|
Every selector lives in a provider configuration file
|
|
(`~/.llm-gateway/providers/<provider>.toml`), never in the Rust source. When a
|
|
web UI changes, edit the TOML: the server picks the change up within a second
|
|
(no restart, no browser relaunch). `llm-gateway selftest <provider>` tells you
|
|
whether the current selectors still work.
|
|
|
|
## Why this file exists
|
|
|
|
Web UIs change without notice. The three failure modes we care about are:
|
|
|
|
| Symptom | Reported as | What to do |
|
|
|---|---|---|
|
|
| The prompt field disappeared | HTTP 500, code `selector_missing` | Update `selectors.input_field` |
|
|
| The send button disappeared | HTTP 500, code `selector_missing` | Update `selectors.send_button` |
|
|
| The answer container moved | empty answer or HTTP 500 | Update `selectors.response_container` |
|
|
| The streaming indicator is wrong | answers truncated or 504 | Update `selectors.streaming_indicator` |
|
|
| The session expired | HTTP 401, code `requires_login` | Run `llm-gateway login <provider>` |
|
|
|
|
A failing capture always leaves a screenshot and a DOM dump in
|
|
`~/.llm-gateway/debug/`, plus one line per failure in `debug/index.jsonl`.
|
|
|
|
## How to update a selector
|
|
|
|
1. Open the provider page in a normal browser, signed in.
|
|
2. Open the developer tools, inspect the element you need.
|
|
3. Prefer stable hooks in this order: `data-testid` > `id` > semantic attributes
|
|
(`aria-label`, `role`, `type`) > a CSS class. Avoid generated class names such
|
|
as `.css-1x2y3z`: they change on every deploy.
|
|
4. Put a comma separated list in the TOML when a UI has several variants
|
|
(`div#prompt-textarea, div.ProseMirror[contenteditable='true']`): the first
|
|
match wins and the fallbacks keep you running through a partial rollout.
|
|
5. Run `llm-gateway selftest <provider>` and keep the TOML only when it reports
|
|
`ok`.
|
|
|
|
## Required and optional selectors
|
|
|
|
| Key | Required | Used for |
|
|
|---|---|---|
|
|
| `input_field` | yes | focusing the editor, verifying the injected text, detecting an empty editor after submit |
|
|
| `send_button` | yes | submitting the prompt (Enter is the fallback) |
|
|
| `response_container` | yes | one node per assistant answer, in document order; the driver reads the n-th node. It must select the **answer text**, not the whole exchange: a container that also holds the action toolbar would capture the word "Copy" as part of the answer |
|
|
| `copy_button` | no | the per-answer "copy" control, which gives the answer as Markdown through the clipboard. It may be inside the answer node or a sibling of it: the driver looks for it there and never elsewhere on the page, so the copy button of a previous answer is never clicked |
|
|
| `streaming_indicator` | recommended | element that only exists while the answer is generated (a stop button) |
|
|
| `file_input` | for images | hidden input fed through `DOM.setFileInputFiles` |
|
|
| `attachment_indicator` | optional | chip shown once an upload has been ingested |
|
|
| `new_chat_button` | optional | starting a brand new conversation without reloading |
|
|
| `login_page_indicator` | recommended | detecting a sign-in page even when the URL does not say so. On chatgpt.com the signed-out landing page exposes `button[data-testid='login']` |
|
|
| `[login] text_patterns` | recommended | lowercase substrings of the visible text that reveal a signed-out page; only consulted when the prompt field is missing, so a chat page that mentions "log in" is never misclassified |
|
|
| `captcha_indicator` | recommended | detecting a bot challenge |
|
|
| `error_banner` | recommended | surfacing the upstream error message to the API client |
|
|
| `model_selector` | reserved | not driven yet: the gateway never changes the model in the UI |
|
|
|
|
## Answer fidelity (Markdown)
|
|
|
|
A web UI renders its answer as rich HTML and `innerText` flattens it: a table
|
|
collapses into loose lines and a code block loses its fences. Once the answer has
|
|
settled the driver therefore reads it a second time, in this order:
|
|
|
|
1. **the clipboard** — it clicks the provider's own copy button and reads the
|
|
clipboard over CDP (`Browser.setPermission` then
|
|
`navigator.clipboard.readText()`). The provider copies exactly the Markdown it
|
|
rendered, so nothing is guessed. Requires `selectors.copy_button`.
|
|
2. **the DOM** — it converts the answer node's `innerHTML` to Markdown
|
|
(headings, lists, quotes, links, images, fenced code, GitHub tables).
|
|
3. **the visible text** — the historical behaviour, used when both fail.
|
|
|
|
`[capture] markdown` chooses between them and a provider may override it in its
|
|
own `[input]` section. Streamed deltas always come from the DOM, so they stay
|
|
plain text; the final answer — the one a non-streaming client receives — is the
|
|
converted one, and the source used is logged as `fidelity`.
|
|
|
|
## Current status
|
|
|
|
| Provider | File | Last validated | Status |
|
|
|---|---|---|---|
|
|
| OpenAI ChatGPT | `providers/chatgpt.toml` | 2026-09-18, signed-in session | `selftest` ok; `fidelity = "clipboard"` |
|
|
| Anthropic Claude | `providers/claude.toml` | 2026-09-18, signed-in session | `selftest` ok; `fidelity = "clipboard"` |
|
|
| DeepSeek | `providers/deepseek.toml` | 2026-09-18, signed-in session | `selftest` ok; **`fidelity = "dom"`**: the copy button has not been observed yet |
|
|
| Local fixture (tests only) | generated by `tests/e2e_browser.rs` | 2026-09-18 | validated: full pipeline, streaming, login, captcha, rate limit, error banner, timeout, missing selector, rewrite, attachment, Markdown fidelity |
|
|
|
|
The Claude and DeepSeek selectors were collected from public sources rather than
|
|
observed on a signed-in session. Treat them as a starting point: the first run of
|
|
`llm-gateway selftest <provider>` tells you whether they still match, and the
|
|
failure modes above tell you which key to fix.
|
|
|
|
The ChatGPT selectors shipped in this repository are the ones observed on the
|
|
public web UI in September 2026. They are a starting point, not a guarantee:
|
|
run the selftest before relying on them.
|
|
|
|
## Upgrading the installed files
|
|
|
|
The provider files under `~/.llm-gateway/providers/` are **copies** made on the
|
|
first run: editing the ones in this repository does not change them. After
|
|
upgrading the binary, refresh them with:
|
|
|
|
```bash
|
|
llm-gateway config init --force # keeps the previous file as <provider>.toml.bak
|
|
```
|
|
|
|
## Localised user interfaces
|
|
|
|
A provider translates its own UI, and the browser locale of the machine does not
|
|
change that: the same ChatGPT account serves a French page under `lang="fr-CA"`
|
|
and an English one under `lang="en-US"`. Every `aria-label` fallback therefore
|
|
only matches one language, which is why the stable hooks come first:
|
|
|
|
| Observed (French UI, September 2026) | Selector to prefer |
|
|
|---|---|
|
|
| `<button type="submit" id="composer-submit-button" aria-label="Envoyer la requête" data-testid="send-button">` | `button[data-testid='send-button']`, `button#composer-submit-button` |
|
|
| the answer toolbar copy control | `button[data-testid='copy-turn-action-button']` |
|
|
| `<textarea placeholder="Message DeepSeek" name="search" rows="2">` | `textarea[placeholder*='Message DeepSeek']` - the English placeholder and the `#chat-input` id of older builds are gone |
|
|
|
|
The `[login] text_patterns` follow the same rule: they are lowercase substrings of
|
|
the page text, so a French sign-in page needs French patterns
|
|
(`"se connecter"`, `"s'inscrire"`) next to the English ones. A page that is not
|
|
recognised as a sign-in page is reported as a missing prompt field instead, which
|
|
is a confusing way to learn that the session expired.
|
|
|
|
## Several accounts of one provider
|
|
|
|
A signed-in account is a browser profile, so driving two ChatGPT accounts means
|
|
two profiles, two sessions and two tab pools. Declare them in the provider file:
|
|
|
|
```toml
|
|
[[accounts]]
|
|
id = "perso" # the short name used in the model string
|
|
email = "[email protected]" # also accepted as the account key
|
|
label = "Personal" # free-form, shown by 'llm-gateway list'
|
|
default = true # used by {"model": "chatgpt"}
|
|
|
|
[[accounts]]
|
|
id = "pro"
|
|
email = "[email protected]"
|
|
```
|
|
|
|
Then sign each one in once and address it in the model string:
|
|
|
|
```bash
|
|
llm-gateway login chatgpt --account perso
|
|
llm-gateway login chatgpt@pro # the same thing
|
|
```
|
|
|
|
```jsonc
|
|
{"model": "chatgpt"} // the default account
|
|
{"model": "chatgpt@pro"} // by id
|
|
{"model": "[email protected]@gmail.com"} // by address
|
|
{"model": "chatgpt@pro/gpt-4o"} // account and model
|
|
```
|
|
|
|
Their profiles live in `profiles/<provider>/accounts/<account>/user-data-dir`.
|
|
A provider that declares no account keeps the historical
|
|
`profiles/<provider>/user-data-dir`, so adding accounts later means signing in
|
|
again for the first one. Each account has its own tab pool, so `max_tabs`
|
|
applies per account.
|
|
|
|
## Adding a new provider
|
|
|
|
1. Copy `providers/chatgpt.toml` to `~/.llm-gateway/providers/<name>.toml`.
|
|
2. Set `provider.name`, `provider.web_url`, `provider.new_conversation_url` and
|
|
`provider.conversation_url_pattern` (one capture group, the web conversation
|
|
id extracted from the page URL).
|
|
3. Fill in the selectors, then `llm-gateway selftest <name>`.
|
|
4. `llm-gateway login <name>` once, and the provider is usable as a model:
|
|
`{"model": "<name>"}` or `{"model": "<name>/<model>"}`.
|
|
|
|
Nothing else is provider specific: the capture loop, the error mapping, the
|
|
conversation store and the API are shared.
|