Files
bruno 7b5aa6fcc9
CI / fmt, clippy and tests (ubuntu-latest) (push) Successful in 21m0s
CI / browser end-to-end (Chromium + fixture) (push) Skipped
CI / fmt, clippy and tests (windows-latest) (push) Canceled after 0s
Add initial llm-bridge project
2026-09-18 18:01:07 -04:00

9.3 KiB

Provider selectors

Every selector lives in a provider configuration file (~/.llm-gateway/providers/<provider>.toml), never in the Rust source. When a web UI changes, edit the TOML: the server picks the change up within a second (no restart, no browser relaunch). llm-gateway selftest <provider> tells you whether the current selectors still work.

Why this file exists

Web UIs change without notice. The three failure modes we care about are:

Symptom Reported as What to do
The prompt field disappeared HTTP 500, code selector_missing Update selectors.input_field
The send button disappeared HTTP 500, code selector_missing Update selectors.send_button
The answer container moved empty answer or HTTP 500 Update selectors.response_container
The streaming indicator is wrong answers truncated or 504 Update selectors.streaming_indicator
The session expired HTTP 401, code requires_login Run llm-gateway login <provider>

A failing capture always leaves a screenshot and a DOM dump in ~/.llm-gateway/debug/, plus one line per failure in debug/index.jsonl.

How to update a selector

  1. Open the provider page in a normal browser, signed in.
  2. Open the developer tools, inspect the element you need.
  3. Prefer stable hooks in this order: data-testid > id > semantic attributes (aria-label, role, type) > a CSS class. Avoid generated class names such as .css-1x2y3z: they change on every deploy.
  4. Put a comma separated list in the TOML when a UI has several variants (div#prompt-textarea, div.ProseMirror[contenteditable='true']): the first match wins and the fallbacks keep you running through a partial rollout.
  5. Run llm-gateway selftest <provider> and keep the TOML only when it reports ok.

Required and optional selectors

Key Required Used for
input_field yes focusing the editor, verifying the injected text, detecting an empty editor after submit
send_button yes submitting the prompt (Enter is the fallback)
response_container yes one node per assistant answer, in document order; the driver reads the n-th node. It must select the answer text, not the whole exchange: a container that also holds the action toolbar would capture the word "Copy" as part of the answer
copy_button no the per-answer "copy" control, which gives the answer as Markdown through the clipboard. It may be inside the answer node or a sibling of it: the driver looks for it there and never elsewhere on the page, so the copy button of a previous answer is never clicked
streaming_indicator recommended element that only exists while the answer is generated (a stop button)
file_input for images hidden input fed through DOM.setFileInputFiles
attachment_indicator optional chip shown once an upload has been ingested
new_chat_button optional starting a brand new conversation without reloading
login_page_indicator recommended detecting a sign-in page even when the URL does not say so. On chatgpt.com the signed-out landing page exposes button[data-testid='login']
[login] text_patterns recommended lowercase substrings of the visible text that reveal a signed-out page; only consulted when the prompt field is missing, so a chat page that mentions "log in" is never misclassified
captcha_indicator recommended detecting a bot challenge
error_banner recommended surfacing the upstream error message to the API client
model_selector reserved not driven yet: the gateway never changes the model in the UI

Answer fidelity (Markdown)

A web UI renders its answer as rich HTML and innerText flattens it: a table collapses into loose lines and a code block loses its fences. Once the answer has settled the driver therefore reads it a second time, in this order:

  1. the clipboard — it clicks the provider's own copy button and reads the clipboard over CDP (Browser.setPermission then navigator.clipboard.readText()). The provider copies exactly the Markdown it rendered, so nothing is guessed. Requires selectors.copy_button.
  2. the DOM — it converts the answer node's innerHTML to Markdown (headings, lists, quotes, links, images, fenced code, GitHub tables).
  3. the visible text — the historical behaviour, used when both fail.

[capture] markdown chooses between them and a provider may override it in its own [input] section. Streamed deltas always come from the DOM, so they stay plain text; the final answer — the one a non-streaming client receives — is the converted one, and the source used is logged as fidelity.

Current status

Provider File Last validated Status
OpenAI ChatGPT providers/chatgpt.toml 2026-09-18, signed-in session selftest ok; fidelity = "clipboard"
Anthropic Claude providers/claude.toml 2026-09-18, signed-in session selftest ok; fidelity = "clipboard"
DeepSeek providers/deepseek.toml 2026-09-18, signed-in session selftest ok; fidelity = "dom": the copy button has not been observed yet
Local fixture (tests only) generated by tests/e2e_browser.rs 2026-09-18 validated: full pipeline, streaming, login, captcha, rate limit, error banner, timeout, missing selector, rewrite, attachment, Markdown fidelity

The Claude and DeepSeek selectors were collected from public sources rather than observed on a signed-in session. Treat them as a starting point: the first run of llm-gateway selftest <provider> tells you whether they still match, and the failure modes above tell you which key to fix.

The ChatGPT selectors shipped in this repository are the ones observed on the public web UI in September 2026. They are a starting point, not a guarantee: run the selftest before relying on them.

Upgrading the installed files

The provider files under ~/.llm-gateway/providers/ are copies made on the first run: editing the ones in this repository does not change them. After upgrading the binary, refresh them with:

llm-gateway config init --force     # keeps the previous file as <provider>.toml.bak

Localised user interfaces

A provider translates its own UI, and the browser locale of the machine does not change that: the same ChatGPT account serves a French page under lang="fr-CA" and an English one under lang="en-US". Every aria-label fallback therefore only matches one language, which is why the stable hooks come first:

Observed (French UI, September 2026) Selector to prefer
<button type="submit" id="composer-submit-button" aria-label="Envoyer la requête" data-testid="send-button"> button[data-testid='send-button'], button#composer-submit-button
the answer toolbar copy control button[data-testid='copy-turn-action-button']
<textarea placeholder="Message DeepSeek" name="search" rows="2"> textarea[placeholder*='Message DeepSeek'] - the English placeholder and the #chat-input id of older builds are gone

The [login] text_patterns follow the same rule: they are lowercase substrings of the page text, so a French sign-in page needs French patterns ("se connecter", "s'inscrire") next to the English ones. A page that is not recognised as a sign-in page is reported as a missing prompt field instead, which is a confusing way to learn that the session expired.

Several accounts of one provider

A signed-in account is a browser profile, so driving two ChatGPT accounts means two profiles, two sessions and two tab pools. Declare them in the provider file:

[[accounts]]
id = "perso"                          # the short name used in the model string
email = "[email protected]"     # also accepted as the account key
label = "Personal"                    # free-form, shown by 'llm-gateway list'
default = true                        # used by {"model": "chatgpt"}

[[accounts]]
id = "pro"
email = "[email protected]"

Then sign each one in once and address it in the model string:

llm-gateway login chatgpt --account perso
llm-gateway login chatgpt@pro                  # the same thing
{"model": "chatgpt"}                                  // the default account
{"model": "chatgpt@pro"}                              // by id
{"model": "[email protected]@gmail.com"}         // by address
{"model": "chatgpt@pro/gpt-4o"}                       // account and model

Their profiles live in profiles/<provider>/accounts/<account>/user-data-dir. A provider that declares no account keeps the historical profiles/<provider>/user-data-dir, so adding accounts later means signing in again for the first one. Each account has its own tab pool, so max_tabs applies per account.

Adding a new provider

  1. Copy providers/chatgpt.toml to ~/.llm-gateway/providers/<name>.toml.
  2. Set provider.name, provider.web_url, provider.new_conversation_url and provider.conversation_url_pattern (one capture group, the web conversation id extracted from the page URL).
  3. Fill in the selectors, then llm-gateway selftest <name>.
  4. llm-gateway login <name> once, and the provider is usable as a model: {"model": "<name>"} or {"model": "<name>/<model>"}.

Nothing else is provider specific: the capture loop, the error mapping, the conversation store and the API are shared.