Compare commits

...
5 Commits
Author SHA1 Message Date
bruno 2c460022f8 test(pdf): helper d'ouverture robuste au vault replie (BUG-060)
CI / lint (push) Successful in 1m37s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 3m33s
CI / build (push) Successful in 59s
CI / e2e (push) Successful in 10m57s
2026-09-17 19:42:43 -04:00
bruno 133644a0ba fix(pdf): affichage des pages du viewer PDF via iframe (BUG-060)
CI / lint (push) Successful in 1m36s
CI / security (push) Successful in 1m3s
CI / test (push) Failing after 3m41s
CI / build (push) Skipped
CI / e2e (push) Skipped
2026-09-17 19:39:48 -04:00
bruno ba0ec3d1fa feat(config): cles des sources connectees et recherche a cle editables depuis la page Configurations (#103)
CI / lint (push) Successful in 1m46s
CI / security (push) Successful in 1m23s
CI / test (push) Successful in 3m42s
CI / build (push) Successful in 58s
CI / e2e (push) Successful in 10m50s
2026-09-17 13:59:26 -04:00
bruno 22e9240e4f fix(ai): les clics ne font plus sauter la conversation au bas (BUG-059); tableaux markdown corrects dans create_pdf (#92) 2026-09-17 13:55:36 -04:00
bruno 6a58a59a11 feat(ai): ecosysteme d'outils phase 2 - recherche a cle, cache/retry, Playwright, crawl, Gitea/GitHub, documents (#92)
CI / lint (push) Successful in 1m37s
CI / security (push) Successful in 1m1s
CI / test (push) Successful in 3m26s
CI / build (push) Successful in 1m44s
CI / e2e (push) Successful in 11m1s
2026-09-17 11:52:03 -04:00
45 changed files with 3176 additions and 43 deletions
+16
View File
@@ -74,3 +74,19 @@ DEEPSEEK_MODEL=deepseek-chat
# Chaîne de repli sans clé (DuckDuckGo puis Bing) si SearXNG ne remonte rien
# OBSIGATE_WEB_FALLBACK=1
# OBSIGATE_WEB_TIMEOUT=10
# Fournisseurs à clé (#92), essayés avant SearXNG — injecter via Infisical en prod
# OBSIGATE_TAVILY_API_KEY=
# OBSIGATE_BRAVE_API_KEY=
# OBSIGATE_SERPAPI_API_KEY=
# OBSIGATE_EXA_API_KEY=
# Ordre des fournisseurs (sinon : clés présentes puis SearXNG puis replis)
# OBSIGATE_WEB_PROVIDERS=brave,searxng
# Réessais réseau (backoff maison) + cache SQLite des résultats web
# OBSIGATE_WEB_RETRY=1
# OBSIGATE_WEB_CACHE_TTL=900 # secondes ; 0 = cache désactivé
# Rendu dynamique (pages SPA) — dépendance optionnelle :
# pip install playwright && playwright install chromium
# ── Assistant IA — sources connectées (Gitea / GitHub) ──
# OBSIGATE_GITEA_URL=https://git.example.net
# OBSIGATE_GITEA_TOKEN=
# OBSIGATE_GITHUB_TOKEN=
+1
View File
@@ -38,6 +38,7 @@ jobs:
- name: Frontend unit tests
run: |
node tests/frontend/unit.test.mjs
node tests/frontend/pdf-viewer.test.mjs
node tests/frontend/forge-completion.test.mjs
- name: Frontend JSDOM tests (PaneManager + Excalidraw + Plugins + AI + SW + Collab + Mobile + Semantic + Desktop + Inline edition)
+93 -1
View File
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
> **En cours de développement** : les changements à venir sont listés dans la section
> [Unreleased](#unreleased). La dernière version livrée est **2.9.1**.
> [Unreleased](#unreleased). La dernière version livrée est **2.11.2**.
---
@@ -14,6 +14,98 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
---
## [2.11.2] — 2026-09-17
### Modifié
- **Tests E2E PDF (BUG-060)** : le helper d'ouverture de fichier de
`tests/e2e/pdf-viewer.spec.js` étend désormais le vault dans l'arborescence s'il est replié
(cas d'une instance avec authentification activée) au lieu de supposer le vault déjà déplié.
---
## [2.11.1] — 2026-09-17
### Corrigé
- **BUG-060 — Viewer PDF : l'affichage des pages ne fonctionnait pas** : au clic sur un fichier
`.pdf`, seule la barre d'outils « PDF — N pages » s'affichait, le corps restant vide. La CSP
durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>`, alors
que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était
bloqué. Le rendu passe désormais par une `<iframe>` (autorisée par `frame-src 'self'`, le stream
`/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` est conservé. Fichiers :
`frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau),
`tests/e2e/pdf-viewer.spec.js` (nouveau, fixture `test_vault/sample-pdf.pdf`).
---
## [2.11.0] — 2026-09-17
### Ajouté
- **#103 — Clés des sources connectées configurables depuis la page Configurations** : nouvelle
section « 🔗 Sources connectées & recherche » (admin) permettant de saisir ses propres jetons
sans toucher aux variables d'environnement : clés de recherche web à clé (Tavily, Brave,
SerpAPI, Exa), URL Gitea + token, token GitHub. Stockage dans `data/api_keys.json` (même
fichier que les clés des fournisseurs IA) via les endpoints
`GET/POST/DELETE /api/config/tool-keys` (GET masque les secrets, URLs en clair ; liste
blanche stricte ; admin requis). Les outils lisent la valeur stockée **en priorité** puis
l'environnement (Infisical) — `backend/tools/secrets.py`. Fichiers :
`backend/tools/{secrets,web,connected}.py`, `backend/main.py`, `frontend/index.html`,
`frontend/js/config.js`, `frontend/locales/{fr,en}.json`, `tests/test_tool_keys.py` (nouveau),
`tests/test_connected_tools.py`.
---
## [2.10.1] — 2026-09-17
### Corrigé
- **BUG-059 — Assistant IA : un clic dans la conversation faisait sauter le texte au bas de la
fenêtre** : dans une conversation ouverte (post ancré en haut), tout clic — lien de fichier,
étapes, sélection de texte — dépinait l'ancre via le gestionnaire `mousedown` et retirait le
padding d'ancre, bornant le scroll à la nouvelle hauteur max (saut au bas). Le dépintage est
désormais réservé aux vrais gestes de scroll : molette, tactile et poignée de scroll
uniquement (`isScrollbarPress`). Fichiers : `frontend/js/bookslm.js`,
`tests/frontend/ai.test.mjs`.
- **#92 — `create_pdf` de l'Assistant : tableaux mal formatés** : l'outil `create_pdf`
(génération PDF via l'assistant) utilisait un rendu reportlab simplifié sans support des
tableaux. Il passe désormais par le même pipeline que le bouton « Télécharger PDF » de la
page document (mistune + plugin `table` + WeasyPrint), avec repli automatique sur le rendu
simple si WeasyPrint n'est pas disponible (GTK absent). Fichiers :
`backend/tools/documents.py`, `tests/test_document_tools.py`.
---
## [2.10.0] — 2026-09-17
### Ajouté
- **#92 — Assistant IA : écosystème d'outils phase 2 (web étendu, sources connectées, documents)** :
- **Recherche web à clé** : fournisseurs optionnels essayés avant SearXNG — Tavily, Brave
Search, SerpAPI (Google) et Exa (`OBSIGATE_TAVILY_API_KEY`, `OBSIGATE_BRAVE_API_KEY`,
`OBSIGATE_SERPAPI_API_KEY`, `OBSIGATE_EXA_API_KEY`), avec ordre personnalisable via
`OBSIGATE_WEB_PROVIDERS`.
- **Transverse** : réessais réseau avec backoff maison (`OBSIGATE_WEB_RETRY`) et cache SQLite
des résultats web avec TTL (`OBSIGATE_WEB_CACHE_TTL`, 900 s par défaut, `OBSIGATE_WEB_CACHE_PATH`).
- **Rendu dynamique** : `fetch_url(render=True)` délègue les pages SPA à un worker Playwright
isolé (dépendance optionnelle, dégradation propre si non installée).
- **Crawl de site** : `crawl_site` (WRITE, confirmation) capture jusqu'à 20 pages d'un même
hôte et enregistre un condensé Markdown dans un vault.
- **Sources connectées** : Gitea (`OBSIGATE_GITEA_URL`/`OBSIGATE_GITEA_TOKEN`) et GitHub
(`OBSIGATE_GITHUB_TOKEN`) — `git_list_repos`, `git_search_issues`, `git_get_file` (READ,
rate-limités, audités). Les drives cloud (Google Drive / OneDrive) restent orientés serveur
MCP externe (#79), conformément à la feuille de route.
- **Production de documents** : `create_xlsx` (openpyxl), `create_docx` (python-docx),
`create_csv`, `create_pdf` (reportlab) — outils WRITE avec confirmation et sauvegarde
dans le vault (backup avant écrasement).
- Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests unitaires mockés
(httpx). Dépendances ajoutées : `openpyxl`, `python-docx`, `reportlab`.
Fichiers : `backend/tools/{webcache,webrender,connected,crawler,documents,web,schemas,labels}.py`,
`backend/services/mutations.py`, `tests/test_{web_cache,web_search_providers,webrender,connected_tools,document_tools,crawler}.py`.
---
## [2.9.1] — 2026-09-17
### Corrigé
+13 -3
View File
@@ -4,7 +4,7 @@
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
[![Version](https://img.shields.io/badge/Version-2.9.1-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.11.2-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -282,6 +282,16 @@ Un compte **admin** connecté voit une icône 🛡️ dans le header : liste, cr
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Autoriser les webhooks vers des adresses privées/boucle | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Taille max des PDF extraits (text indexation) | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | Timeout extraction PDF (secondes) | `30` |
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Fournisseurs de recherche web à clé (essayés avant SearXNG) | — |
| `OBSIGATE_WEB_PROVIDERS` | Ordre des fournisseurs de recherche (ex. `brave,searxng`) | — |
| `OBSIGATE_WEB_RETRY` | Réessais réseau des outils web (backoff maison) | `1` |
| `OBSIGATE_WEB_CACHE_TTL` | Durée du cache SQLite des résultats web (secondes, `0` = off) | `900` |
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Source connectée Gitea (outil `git_list_repos`…) | — |
| `OBSIGATE_GITHUB_TOKEN` | Jeton GitHub (outil `git_list_repos`…) | — |
> Ces clés peuvent aussi être saisies **depuis l'interface** (menu → Configurations →
> « Sources connectées & recherche ») : la valeur saisie est stockée dans `data/api_keys.json`
> et prime sur la variable d'environnement.
### Volume pour la persistance
@@ -916,8 +926,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
## 📝 Changelog
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.9.1).
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.11.2).
---
*Projet : ObsiGate | Version : 2.9.1 | Dernière mise à jour : Juin 2026*
*Projet : ObsiGate | Version : 2.11.2 | Dernière mise à jour : Juin 2026*
+13 -3
View File
@@ -2,7 +2,7 @@
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
[![Version](https://img.shields.io/badge/Version-2.9.1-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.11.2-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -320,6 +320,16 @@ When an **admin** account is logged in, a 🛡️ icon appears in the header. Cl
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Allow webhooks to private/loopback addresses | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Max PDF size for text extraction | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | PDF extraction timeout (seconds) | `30` |
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Keyed web-search providers (tried before SearXNG) | — |
| `OBSIGATE_WEB_PROVIDERS` | Search provider order (e.g. `brave,searxng`) | — |
| `OBSIGATE_WEB_RETRY` | Web tools network retries (house-made backoff) | `1` |
| `OBSIGATE_WEB_CACHE_TTL` | SQLite cache TTL for web results (seconds, `0` = off) | `900` |
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Gitea connected source (`git_list_repos`…) | — |
| `OBSIGATE_GITHUB_TOKEN` | GitHub token (`git_list_repos`…) | — |
> These keys can also be entered **from the UI** (menu → Configurations →
> "Connected sources & search"): the stored value goes to `data/api_keys.json`
> and takes precedence over the environment variable.
>All these variables are documented in `.env.example`.
@@ -1085,8 +1095,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
## 📝 Changelog
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.9.1).
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.11.2).
---
*Project: ObsiGate | Version: 2.9.1 | Last updated: May 2026*
*Project: ObsiGate | Version: 2.11.2 | Last updated: May 2026*
+1 -1
View File
@@ -1 +1 @@
2.9.1
2.11.2
+62
View File
@@ -3679,6 +3679,68 @@ async def api_delete_ai_key(provider_env: str, current_user=Depends(require_admi
return {"status": "deleted", "key": key_name}
# ---------------------------------------------------------------------------
# Tool & connected-source keys (#103) — same store as the AI provider keys
# ---------------------------------------------------------------------------
from backend.tools.secrets import (
TOOL_KEY_NAMES as _TOOL_KEY_NAMES,
)
from backend.tools.secrets import (
delete_tool_key as _delete_tool_key,
)
from backend.tools.secrets import (
get_tool_key as _get_tool_key,
)
from backend.tools.secrets import (
mask_value as _mask_tool_value,
)
from backend.tools.secrets import (
set_tool_key as _set_tool_key,
)
@app.get("/api/config/tool-keys", response_model=AIKeysResponse)
async def api_get_tool_keys(current_user=Depends(require_admin)):
"""Return tool/connected-source configuration (tokens masked, URLs clear)."""
masked = {}
for name in _TOOL_KEY_NAMES:
masked[name] = _mask_tool_value(name, _get_tool_key(name))
return masked
@app.post("/api/config/tool-keys", response_model=StatusResponse)
async def api_set_tool_keys(body: dict = Body(...), current_user=Depends(require_admin)):
"""Save tool/connected-source keys.
Only whitelisted names (``backend.tools.secrets.TOOL_KEY_NAMES``) are
accepted: Tavily/Brave/SerpAPI/Exa API keys, Gitea URL + token, GitHub
token. Empty values delete the stored entry.
"""
updated = []
for name, value in body.items():
if name not in _TOOL_KEY_NAMES:
raise HTTPException(status_code=400, detail=f"Clé inconnue: {name}")
if value is not None and not isinstance(value, str):
raise HTTPException(status_code=400, detail=f"Type invalide pour {name}")
_set_tool_key(name, value or "")
updated.append(name)
logger.info(f"Tool keys updated: {updated}")
return {"status": "ok"}
@app.delete("/api/config/tool-keys/{name}", response_model=AIKeyDeleteResponse)
async def api_delete_tool_key(name: str, current_user=Depends(require_admin)):
"""Delete a stored tool key (the environment fallback still applies)."""
key_name = name.upper()
try:
existed = _delete_tool_key(key_name)
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
logger.info(f"Tool key deleted: {key_name} (existed={existed})")
return {"status": "deleted", "key": key_name}
@app.post("/api/config/ai-keys/test", response_model=AITestResponse)
async def api_test_ai_keys(current_user=Depends(require_admin)):
"""Test which AI providers are configured.
+3
View File
@@ -20,3 +20,6 @@ psutil>=5.9
pywebpush>=2.3.0
mcp==1.9.4
sse-starlette==2.1.3
openpyxl>=3.1
python-docx>=1.1
reportlab>=4.0
+9 -3
View File
@@ -44,7 +44,7 @@ def _ensure_writable(root: Path) -> None:
raise ServiceError("Vault is read-only", code="read_only", status=403)
def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
def _validate_extension(file_path: Path, *, allow_images: bool = False, allow_docs: bool = False) -> None:
"""Reject unsupported file extensions (400)."""
from backend.indexer import SUPPORTED_EXTENSIONS
@@ -53,6 +53,9 @@ def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
if allow_images:
from backend.attachment_indexer import IMAGE_EXTENSIONS
allowed = allowed | IMAGE_EXTENSIONS
if allow_docs:
# Office documents produced by the AI tool layer (#92).
allowed = allowed | {".xlsx", ".docx"}
if ext not in allowed and file_path.name.lower() not in ("dockerfile", "makefile"):
raise ServiceError(
@@ -715,17 +718,20 @@ def save_raw_file(
content: bytes,
*,
overwrite: bool = True,
allow_docs: bool = False,
) -> dict[str, Any]:
"""Save a binary or text file to a vault (e.g. from upload / drag-and-drop).
Creates parent directories automatically and safely validates the path.
Supports supported text extensions, images and Excalidraw files.
Supports supported text extensions, images, Excalidraw files and — with
``allow_docs`` — Office documents (.xlsx/.docx) produced by the AI tools.
Args:
vault_name: Name of the vault.
path: Vault-relative path.
content: Raw bytes to write.
overwrite: When True, replace existing files (with backup).
allow_docs: Also accept .xlsx/.docx extensions (AI document tools).
Returns:
Dict with ``success``, ``vault``, ``path``, and ``size``.
@@ -733,7 +739,7 @@ def save_raw_file(
root = get_vault_root(vault_name)
_ensure_writable(root)
file_path = resolve_safe_path(root, path)
_validate_extension(file_path, allow_images=True)
_validate_extension(file_path, allow_images=True, allow_docs=allow_docs)
rel_path = _rel(root, file_path)
+3
View File
@@ -9,6 +9,9 @@ Note: ObsiGate uses implicit namespace packages (no tracked ``__init__.py``,
which ``.gitignore`` excludes via ``_*.py``), hence this explicit facade.
"""
from backend.tools import connected as _connected # noqa: F401 (registers connected-source tools)
from backend.tools import crawler as _crawler # noqa: F401 (registers the site crawler)
from backend.tools import documents as _documents # noqa: F401 (registers document tools)
from backend.tools import service as _service # noqa: F401 (registers tools)
from backend.tools import web as _web # noqa: F401 (registers web tools)
from backend.tools.context import (
+241
View File
@@ -0,0 +1,241 @@
"""Connected sources — Gitea & GitHub repositories (phase 2 #92).
The assistant can query the source-hosting platforms the project actually
uses (ObsiGate is hosted on Gitea): repositories, issues/pull requests and
repository files. Everything is READ-risk, rate-limited through the shared
registry and audited.
Configuration (environment — injected by Infisical in production, never
hard-coded):
* ``OBSIGATE_GITEA_URL`` — base URL of the self-hosted instance (e.g.
``https://git.example.net``); the ``gitea`` provider is only available when
this variable is set. Admin-controlled, so the SSRF guard does not apply
(unlike user-supplied URLs). Both the URL and the tokens can also be set
from the configuration page (stored in ``data/api_keys.json``, #103) —
the stored value takes precedence over the environment.
* ``OBSIGATE_GITEA_TOKEN`` — optional personal access token (private repos).
* ``OBSIGATE_GITHUB_TOKEN`` — optional token (raises the API rate limits and
unlocks private repositories).
Cloud drives (Google Drive / OneDrive) deliberately stay out of the core:
per the documented roadmap they are best served by an *external MCP server*
(#79) so the OAuth surface remains outside ObsiGate.
"""
from __future__ import annotations
import base64
import binascii
import logging
from typing import Any
import httpx
from backend.tools.context import ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import GitGetFileInput, GitProviderInput, GitSearchIssuesInput
from backend.tools.secrets import get_tool_key
logger = logging.getLogger("obsigate.tools.connected")
TIMEOUT = 10.0
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
MAX_FILE_BYTES = 300_000
GITHUB_API = "https://api.github.com"
def _provider_base(provider: str) -> tuple[str, str]:
"""Return (base_url, auth_header_value) for the requested provider."""
if provider == "gitea":
base = get_tool_key("OBSIGATE_GITEA_URL").rstrip("/")
if not base:
raise ToolError(
"Source Gitea non configurée (OBSIGATE_GITEA_URL absente).",
code="provider_not_configured",
)
token = get_tool_key("OBSIGATE_GITEA_TOKEN")
return base, f"token {token}" if token else ""
if provider == "github":
token = get_tool_key("OBSIGATE_GITHUB_TOKEN")
return GITHUB_API, f"Bearer {token}" if token else ""
raise ToolError(
f"Fournisseur inconnu : {provider} ('gitea' ou 'github')",
code="invalid_arguments",
)
def _headers(auth: str) -> dict[str, str]:
headers = {"User-Agent": USER_AGENT, "Accept": "application/json"}
if auth:
headers["Authorization"] = auth
return headers
def _request(method: str, url: str, auth: str, **kwargs: Any) -> httpx.Response:
try:
resp = httpx.request(
method, url, headers=_headers(auth), timeout=TIMEOUT, follow_redirects=False,
**kwargs,
)
except httpx.HTTPError as e:
logger.warning("connected source request failed %s: %s", url, e)
raise ToolError(
"Source connectée momentanément indisponible.",
code="connected_source_unavailable",
) from e
if resp.status_code in (401, 403):
raise ToolError(
"Accès refusé par la source connectée (jeton manquant ou expiré).",
code="permission_denied",
)
if resp.status_code == 404:
raise ToolError("Ressource introuvable sur la source connectée.", code="not_found")
resp.raise_for_status()
return resp
def _normalize_repo(item: dict[str, Any]) -> dict[str, Any]:
return {
"name": item.get("name") or "",
"full_name": item.get("full_name") or "",
"url": item.get("html_url") or item.get("clone_url") or "",
"description": item.get("description") or "",
"updated": item.get("updated_at") or "",
"private": bool(item.get("private", False)),
}
@tool(
name="git_list_repos",
description=(
"List repositories on the connected Gitea instance or GitHub account "
"(name, url, description, last update). Use when the user asks about "
"their code projects."
),
input_model=GitProviderInput,
risk=ToolRisk.READ,
)
def git_list_repos(ctx, params: GitProviderInput) -> dict[str, Any]:
"""Query the configured source and return normalized repositories."""
base, auth = _provider_base(params.provider)
if params.provider == "gitea":
url = base + "/api/v1/repos/search"
query: dict[str, Any] = {"limit": params.limit}
if params.repo:
query["q"] = params.repo
resp = _request("GET", url, auth, params=query)
items = resp.json().get("data") or []
else:
if params.repo:
url = GITHUB_API + f"/repos/{params.repo.strip('/')}"
items = [_request("GET", url, auth).json()]
else:
resp = _request(
"GET", GITHUB_API + "/user/repos",
auth, params={"per_page": params.limit, "sort": "updated"},
)
items = resp.json()
repos = [_normalize_repo(item) for item in items if isinstance(item, dict)]
return {"provider": params.provider, "count": len(repos), "repos": repos}
@tool(
name="git_search_issues",
description=(
"Search issues and pull requests on the connected Gitea instance or "
"GitHub (title/body keywords, optional repository scope, open/closed)."
),
input_model=GitSearchIssuesInput,
risk=ToolRisk.READ,
)
def git_search_issues(ctx, params: GitSearchIssuesInput) -> dict[str, Any]:
"""Query issues (and PRs) from the configured source."""
base, auth = _provider_base(params.provider)
state = params.state if params.state in ("open", "closed") else "open"
if params.provider == "gitea":
if params.repo:
url = base + f"/api/v1/repos/{params.repo.strip('/')}/issues"
query: dict[str, Any] = {"state": state, "limit": params.limit, "q": params.query}
resp = _request("GET", url, auth, params=query)
items = resp.json()
else:
url = base + "/api/v1/repos/issues/search"
resp = _request("GET", url, auth, params={
"q": params.query, "state": state, "limit": params.limit,
})
items = resp.json()
else:
clause = f"{params.query} is:issue is:{state}"
if params.repo:
clause += f" repo:{params.repo.strip('/')}"
resp = _request(
"GET", GITHUB_API + "/search/issues", auth,
params={"q": clause, "per_page": params.limit},
)
items = (resp.json().get("items") or [])
issues = [
{
"id": item.get("number") or item.get("id") or "",
"title": (item.get("title") or "")[:300],
"url": item.get("html_url") or "",
"state": item.get("state") or "",
"pull_request": bool(item.get("pull_request")),
}
for item in (items if isinstance(items, list) else [])
if isinstance(item, dict)
]
return {
"provider": params.provider,
"query": params.query,
"count": len(issues),
"issues": issues,
}
@tool(
name="git_get_file",
description=(
"Read a file's content from a connected Gitea or GitHub repository "
"(source code, docs, config). Text/JSON only, size-capped."
),
input_model=GitGetFileInput,
risk=ToolRisk.READ,
)
def git_get_file(ctx, params: GitGetFileInput) -> dict[str, Any]:
"""Fetch one repository file and return its decoded text content."""
base, auth = _provider_base(params.provider)
repo = params.repo.strip("/")
path = params.path.strip("/")
if not repo or not path:
raise ToolError(
"'repo' (owner/nom) et 'path' sont obligatoires", code="invalid_arguments"
)
if params.provider == "gitea":
url = base + f"/api/v1/repos/{repo}/contents/{path}"
else:
url = GITHUB_API + f"/repos/{repo}/contents/{path}"
if params.ref:
url += f"?ref={params.ref}"
resp = _request("GET", url, auth)
data = resp.json()
encoded = data.get("content") or ""
if (data.get("encoding") or "") == "base64" and encoded:
try:
content = base64.b64decode(encoded).decode("utf-8", errors="replace")
except (ValueError, binascii.Error) as e:
raise ToolError(
"Contenu du fichier illisible (encodage inattendu).",
code="file_decode_error",
) from e
else:
content = encoded
truncated = len(content) > MAX_FILE_BYTES
return {
"provider": params.provider,
"repo": repo,
"path": data.get("path") or path,
"size": data.get("size") or len(content),
"content": content[:MAX_FILE_BYTES],
"truncated": truncated,
}
+197
View File
@@ -0,0 +1,197 @@
"""Multi-page site crawl — ``crawl_site`` (phase 2 #92, WRITE + confirmation).
The assistant can digest a small public site (documentation, docs portal) and
store a Markdown summary inside a vault: one section per page, title, URL and
readable text. The crawl is bounded and same-host only:
* max 20 pages (``max_pages``), same hostname, breadth-first from the entry URL;
* SSRF guard on every URL (scheme + private-address rejection), size caps;
* no third-party crawler dependency (scrapy deliberately avoided — a bounded
httpx BFS keeps the surface small and the runtime predictable; the task is
executed as a single background-style tool run instead of a web request
pipeline).
Risk is WRITE: the digest is written into a vault, so the two-step
confirmation applies (Apply card in the UI, propose/apply over MCP).
"""
from __future__ import annotations
import logging
import re
import time
from typing import Any
from urllib.parse import urljoin, urlparse
import httpx
from backend.tools.context import ToolContext, ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import CrawlSiteInput
from backend.tools.web import (
USER_AGENT,
_assert_public_http_url,
_html_to_text,
_response_text,
)
logger = logging.getLogger("obsigate.tools.crawler")
MAX_PAGE_BYTES = 800_000
MAX_TOTAL_BYTES = 6_000_000
MAX_TEXT_PER_PAGE = 12_000
PAGE_TIMEOUT = 10.0
_LINK_RE = re.compile(r'<a[^>]*href="([^"#]+)"', re.IGNORECASE)
_TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
def _same_host(url: str, host: str) -> bool:
return (urlparse(url).hostname or "") == host
def _extract_links(raw: str, base_url: str) -> list[str]:
import html as html_lib
links: list[str] = []
for match in _LINK_RE.finditer(raw):
href = html_lib.unescape(match.group(1)).strip()
if not href or href.lower().startswith(("javascript:", "mailto:", "tel:")):
continue
absolute = urljoin(base_url, href)
if absolute.lower().endswith((".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp", ".pdf", ".zip")):
continue
links.append(absolute.split("#", 1)[0])
return links
def _fetch_page(url: str) -> tuple[str, str]:
"""Fetch one page (SSRF-guarded, manual redirects) → (title, text)."""
current = _assert_public_http_url(url)
resp = None
for _hop in range(5):
resp = httpx.get(
current,
headers={"User-Agent": USER_AGENT, "Accept": "text/html,*/*"},
timeout=PAGE_TIMEOUT,
follow_redirects=False,
)
if resp.status_code in (301, 302, 303, 307, 308):
location = resp.headers.get("location") or ""
if not location:
break
current = _assert_public_http_url(str(httpx.URL(current).join(location)))
continue
break
assert resp is not None
resp.raise_for_status()
ctype = (resp.headers.get("content-type") or "").lower()
if "html" not in ctype and "text" not in ctype:
raise ToolError(
f"Type de contenu non pris en charge: {ctype.split(';')[0] or 'inconnu'}",
code="unsupported_content_type",
)
raw = (resp.content[:MAX_PAGE_BYTES]).decode(resp.encoding or "utf-8", errors="replace")
title_match = _TITLE_RE.search(raw)
import html as html_lib
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
return title, _html_to_text(raw)[:MAX_TEXT_PER_PAGE]
@tool(
name="crawl_site",
description=(
"Crawl a small public site (same-host only, max 20 pages) starting at "
"a URL and save a Markdown digest (title, url, readable text per page) "
"into a vault. Use to capture an online documentation for offline use."
),
input_model=CrawlSiteInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def crawl_site(ctx: ToolContext, params: CrawlSiteInput) -> dict[str, Any]:
"""Bounded BFS crawl; writes the digest file and returns a summary."""
from backend.services.errors import ServiceError
from backend.services.mutations import save_raw_file
start = _assert_public_http_url(params.url.strip())
host = urlparse(start).hostname or ""
if not host:
raise ToolError("URL sans hôte", code="invalid_url")
queue: list[str] = [start]
seen: set[str] = {start}
pages: list[dict[str, Any]] = []
total_bytes = 0
failures: list[str] = []
while queue and len(pages) < params.max_pages and total_bytes < MAX_TOTAL_BYTES:
url = queue.pop(0)
try:
title, text = _fetch_page(url)
except ToolError as e:
failures.append(url)
logger.warning("crawl_site page failed %s: %s", url, e.code)
continue
except httpx.HTTPError as e:
failures.append(url)
logger.warning("crawl_site page failed %s: %s", url, e)
continue
pages.append({"url": url, "title": title, "text": text})
total_bytes += len(text)
if len(pages) >= params.max_pages:
break
try:
raw_resp = httpx.get(
url, headers={"User-Agent": USER_AGENT}, timeout=PAGE_TIMEOUT,
follow_redirects=False,
)
raw = _response_text(raw_resp)
except (httpx.HTTPError, ValueError):
continue
for link in _extract_links(raw, url):
if len(pages) + len(queue) >= params.max_pages:
break
if link in seen or not _same_host(link, host):
continue
try:
_assert_public_http_url(link)
except ToolError:
continue
seen.add(link)
queue.append(link)
if not pages:
raise ToolError(
"Aucune page n'a pu être récupérée pour ce site.",
code="crawl_failed",
)
lines = [
f"# Crawl de {host}",
"",
f"> {len(pages)} page(s) capturée(s) depuis {start} — {time.strftime('%Y-%m-%d %H:%M')}",
"",
]
for page in pages:
lines.append(f"## {page['title'] or page['url']}")
lines.append("")
lines.append(f"Source : {page['url']}")
lines.append("")
lines.append(page["text"])
lines.append("")
digest = "\n".join(lines).encode("utf-8")
try:
saved = save_raw_file(
params.vault, params.path, digest, overwrite=True, allow_docs=False
)
except ServiceError as e:
raise ToolError(e.message, code=e.code, details=e.details) from e
return {
"url": start,
"vault": params.vault,
"path": saved.get("path", params.path),
"pages": len(pages),
"failed": failures[:10],
"size": saved.get("size", len(digest)),
}
+229
View File
@@ -0,0 +1,229 @@
"""Document-production tools (phase 2 #92) — WRITE, confirmation required.
The assistant can generate real files inside a vault:
* ``create_xlsx`` — spreadsheet (openpyxl);
* ``create_docx`` — Word document (python-docx);
* ``create_csv`` — CSV (stdlib);
* ``create_pdf`` — PDF (reportlab, from markdown-ish content).
Every tool is ``WRITE`` (two-step confirm in the UI / propose-apply over MCP),
vault-scoped through ``requires_vault`` and saved via the shared mutation
service (path safety, read-only check, backup on overwrite).
"""
from __future__ import annotations
import csv as csv_lib
import io
import logging
import re
from typing import Any
from xml.sax import saxutils
from backend.services.errors import ServiceError
from backend.services.mutations import save_raw_file
from backend.tools.context import ToolContext, ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import CsvInput, DocxInput, PdfInput, SpreadsheetInput
logger = logging.getLogger("obsigate.tools.documents")
MAX_PDF_CHARS = 200_000
MAX_ROWS = 5_000
def _save(vault: str, path: str, content: bytes, overwrite: bool) -> dict[str, Any]:
"""Shared save helper (maps ServiceError to ToolError)."""
try:
return save_raw_file(vault, path, content, overwrite=overwrite, allow_docs=True)
except ServiceError as e:
raise ToolError(e.message, code=e.code, details=e.details) from e
def _check_rows(rows: list[list[Any]]) -> None:
if not rows:
raise ToolError("Aucune ligne fournie", code="invalid_arguments")
if len(rows) > MAX_ROWS:
raise ToolError(
f"Trop de lignes ({len(rows)} > {MAX_ROWS})", code="invalid_arguments"
)
def _check_extension(path: str, expected: str) -> str:
"""Enforce the document extension; return the normalized path."""
path = (path or "").strip()
if not path.lower().endswith(expected):
raise ToolError(
f"Extension attendue : {expected}", code="invalid_arguments"
)
return path
@tool(
name="create_xlsx",
description=(
"Create an .xlsx spreadsheet in a vault from rows of cell values "
"(first row = header). Use for tables, budgets, checklists the user "
"asked to turn into an Excel file."
),
input_model=SpreadsheetInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_xlsx(ctx: ToolContext, params: SpreadsheetInput) -> dict[str, Any]:
"""Build the workbook with openpyxl and save it into the vault."""
from openpyxl import Workbook
_check_rows(params.rows)
path = _check_extension(params.path, ".xlsx")
wb = Workbook()
ws = wb.active
ws.title = params.sheet_name[:31] or "Feuille1"
for row in params.rows:
ws.append(list(row))
buffer = io.BytesIO()
wb.save(buffer)
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
@tool(
name="create_docx",
description=(
"Create a .docx Word document in a vault from an optional title and "
"ordered paragraphs. Use for letters, reports, structured drafts."
),
input_model=DocxInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_docx(ctx: ToolContext, params: DocxInput) -> dict[str, Any]:
"""Build the document with python-docx and save it into the vault."""
from docx import Document
if not params.paragraphs:
raise ToolError("Aucun paragraphe fourni", code="invalid_arguments")
path = _check_extension(params.path, ".docx")
doc = Document()
if params.title.strip():
doc.add_heading(params.title.strip(), level=1)
for paragraph in params.paragraphs:
doc.add_paragraph(paragraph)
buffer = io.BytesIO()
doc.save(buffer)
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
@tool(
name="create_csv",
description=(
"Create a .csv file in a vault from rows of cell values (first row = "
"header). Use for flat data exports, simple tables."
),
input_model=CsvInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_csv(ctx: ToolContext, params: CsvInput) -> dict[str, Any]:
"""Serialize the rows and save the CSV into the vault."""
_check_rows(params.rows)
path = _check_extension(params.path, ".csv")
delimiter = params.delimiter if params.delimiter in (",", ";", "\t") else ","
buffer = io.StringIO()
writer = csv_lib.writer(buffer, delimiter=delimiter, lineterminator="\n")
writer.writerows(params.rows)
return _save(params.vault, path, buffer.getvalue().encode("utf-8"), params.overwrite)
_HEADING_RE = re.compile(r"^(#{1,6})\s+(.*)$")
def _markdown_to_flowables(content: str) -> list[tuple[str, str]]:
"""Split markdown-ish content into (style, text) blocks for reportlab."""
blocks: list[tuple[str, str]] = []
for raw_line in content.splitlines():
line = raw_line.rstrip()
if not line.strip():
continue
heading = _HEADING_RE.match(line)
if heading:
blocks.append((f"H{min(3, len(heading.group(1)))}", heading.group(2).strip()))
else:
blocks.append(("P", line.strip()))
return blocks
def _render_markdown_pdf(content: str, title: str) -> bytes | None:
"""Render markdown → HTML → PDF through the document-page pipeline.
Uses the same stack as the « Download PDF » button of the document viewer
(mistune with the table plugin + WeasyPrint print CSS), so tables, code
blocks and lists are laid out correctly. Returns ``None`` when WeasyPrint
is not importable (missing GTK on some hosts) so the caller can fall back
to the simplified reportlab renderer.
"""
try:
import mistune
from backend.pdf_export import build_pdf_html, generate_pdf
renderer = mistune.create_markdown(
escape=False,
plugins=["table", "strikethrough", "footnotes", "task_lists"],
)
html = renderer(content)
return generate_pdf(build_pdf_html(html, title), title)
except Exception as e:
# WeasyPrint loads GTK lazily: a missing native library can surface at
# import OR render time. Fall back to the simple renderer either way.
logger.warning("WeasyPrint pipeline unavailable for create_pdf: %s", e)
return None
def _render_reportlab_pdf(content: str, title: str) -> bytes:
"""Fallback renderer (no WeasyPrint): headings + paragraphs, no tables."""
from reportlab.lib.pagesizes import A4
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.platypus import Paragraph, SimpleDocTemplate, Spacer
styles = getSampleStyleSheet()
style_map = {
"P": styles["BodyText"],
"H1": styles["Heading1"],
"H2": styles["Heading2"],
"H3": styles["Heading3"],
}
buffer = io.BytesIO()
doc = SimpleDocTemplate(buffer, pagesize=A4, title=title[:200])
story: list[Any] = [Paragraph(saxutils.escape(title[:300]), styles["Title"])]
for style, line in _markdown_to_flowables(content):
story.append(Spacer(1, 4))
story.append(Paragraph(saxutils.escape(line), style_map[style]))
doc.build(story)
return buffer.getvalue()
@tool(
name="create_pdf",
description=(
"Create a .pdf document in a vault from markdown content (headings, "
"paragraphs, tables, code blocks, lists). Use for printable "
"deliverables; tables are laid out like the document-page PDF export."
),
input_model=PdfInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_pdf(ctx: ToolContext, params: PdfInput) -> dict[str, Any]:
"""Render the content and save the PDF into the vault.
Primary path: mistune (tables) + WeasyPrint — identical to the viewer's
« Download PDF » export. Fallback (WeasyPrint unavailable): simplified
reportlab layout without tables.
"""
path = _check_extension(params.path, ".pdf")
content = params.content[:MAX_PDF_CHARS]
pdf_bytes = _render_markdown_pdf(content, params.title[:300])
if pdf_bytes is None:
pdf_bytes = _render_reportlab_pdf(content, params.title[:300])
return _save(params.vault, path, pdf_bytes, params.overwrite)
+8
View File
@@ -47,6 +47,14 @@ _STEP_LABELS: dict[str, tuple[str, str | None]] = {
"restore_backup": ("backup_restore", "path"),
"web_search": ("web_search", "query"),
"fetch_url": ("fetch_url", "url"),
"crawl_site": ("crawl", "url"),
"git_list_repos": ("git_repos", "provider"),
"git_search_issues": ("git_issues", "query"),
"git_get_file": ("git_file", "path"),
"create_xlsx": ("xlsx_create", "path"),
"create_docx": ("docx_create", "path"),
"create_csv": ("csv_create", "path"),
"create_pdf": ("pdf_create", "path"),
}
GENERIC_KEY = "generic"
+84
View File
@@ -251,6 +251,90 @@ class FetchUrlInput(BaseModel):
"""Fetch one public web page and return its readable text."""
url: str = Field(..., description="Absolute http(s) URL of a public page")
render: bool = Field(
False,
description="Render JavaScript with the optional Playwright worker (dynamic SPA pages)",
)
class CrawlSiteInput(BaseModel):
"""Crawl a small public site (same-host only) and save a digest into a vault."""
url: str = Field(..., description="Absolute http(s) URL where the crawl starts")
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the digest file to write (.md)")
max_pages: int = Field(5, ge=1, le=20, description="Maximum number of pages to crawl")
class GitProviderInput(BaseModel):
"""Base fields for connected-source tools (Gitea / GitHub)."""
provider: str = Field(..., description="'gitea' (OBSIGATE_GITEA_URL) or 'github'")
repo: str = Field("", description="Optional 'owner/name' repository filter")
limit: int = Field(20, ge=1, le=50, description="Maximum number of entries")
class GitSearchIssuesInput(BaseModel):
"""Search issues/pull requests on a connected Gitea or GitHub instance."""
provider: str = Field(..., description="'gitea' or 'github'")
query: str = Field(..., min_length=1, description="Search keywords")
repo: str = Field("", description="Optional 'owner/name' scope (empty = instance-wide)")
state: str = Field("open", description="'open' or 'closed'")
limit: int = Field(10, ge=1, le=20, description="Maximum number of issues")
class GitGetFileInput(BaseModel):
"""Read a file from a connected Gitea or GitHub repository."""
provider: str = Field(..., description="'gitea' or 'github'")
repo: str = Field(..., description="'owner/name' repository")
path: str = Field(..., description="Repository-relative file path")
ref: str = Field("", description="Optional branch/tag/commit (empty = default branch)")
class SpreadsheetInput(BaseModel):
"""Create an .xlsx spreadsheet in a vault from rows of cells."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.xlsx)")
rows: list[list[str | int | float | bool | None]] = Field(
..., description="Rows of cell values (first row = header)"
)
sheet_name: str = Field("Feuille1", description="Worksheet name")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class DocxInput(BaseModel):
"""Create a .docx Word document in a vault from paragraphs."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.docx)")
title: str = Field("", description="Optional document title (heading 1)")
paragraphs: list[str] = Field(..., description="Paragraph texts, in order")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class CsvInput(BaseModel):
"""Create a .csv file in a vault from rows of cells."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.csv)")
rows: list[list[str | int | float | bool | None]] = Field(
..., description="Rows of cell values (first row = header)"
)
delimiter: str = Field(",", description="Field separator (',' ';' '\\t')")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class PdfInput(BaseModel):
"""Create a .pdf document in a vault from markdown-ish content."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.pdf)")
title: str = Field("Document", description="Document title")
content: str = Field(..., description="Content (headings with #/##, then paragraphs)")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class ToolResult(BaseModel):
+109
View File
@@ -0,0 +1,109 @@
"""Tool-layer secrets — user-configured tokens & API keys (#103).
The connected-source (Gitea / GitHub) and keyed web-search (Tavily, Brave,
SerpAPI, Exa) tools read their credentials through this module instead of
``os.environ`` directly. The value comes from the store the user edits in the
configuration page (``data/api_keys.json`` — the same file the AI provider
keys use) first, then falls back to the environment (Infisical-injected in
production). Nothing is ever hard-coded and no tool result carries a secret
(the registry redacts payloads).
Allowed names are whitelisted: only the variables below can be stored or
deleted from the configuration page.
"""
from __future__ import annotations
import json
import logging
import os
from pathlib import Path
logger = logging.getLogger("obsigate.tools.secrets")
# Whitelisted configuration names (config page « Sources connectées & recherche »).
TOOL_KEY_NAMES: tuple[str, ...] = (
"OBSIGATE_TAVILY_API_KEY",
"OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY",
"OBSIGATE_EXA_API_KEY",
"OBSIGATE_GITEA_URL",
"OBSIGATE_GITEA_TOKEN",
"OBSIGATE_GITHUB_TOKEN",
)
_SECRET_MARKERS = ("API_KEY", "TOKEN")
def _keys_file() -> Path:
base = os.environ.get("OBSIGATE_DATA_DIR", "data")
return Path(base) / "api_keys.json"
def _read_keys() -> dict:
path = _keys_file()
if not path.exists():
return {}
try:
data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, ValueError) as e:
logger.warning("tool key store unreadable (%s): %s", path, e)
return {}
return data if isinstance(data, dict) else {}
def _write_keys(data: dict) -> None:
path = _keys_file()
path.parent.mkdir(parents=True, exist_ok=True)
tmp = path.with_suffix(".tmp")
tmp.write_text(json.dumps(data, indent=2), encoding="utf-8")
tmp.replace(path)
def is_secret_name(name: str) -> bool:
"""True for API keys / tokens (masked in API responses); URLs are clear."""
return any(marker in name for marker in _SECRET_MARKERS)
def mask_value(name: str, value: str) -> str:
"""Mask a secret for display; non-secret values (URLs) are returned as-is."""
if not value:
return ""
if not is_secret_name(name):
return value
return value[:4] + "..." + value[-4:] if len(value) > 8 else "***"
def get_tool_key(name: str) -> str:
"""Stored (configuration page) value first, then environment fallback."""
if name not in TOOL_KEY_NAMES:
return os.environ.get(name, "").strip()
stored = _read_keys().get(name)
if isinstance(stored, str) and stored.strip():
return stored.strip()
return os.environ.get(name, "").strip()
def set_tool_key(name: str, value: str) -> None:
"""Persist one whitelisted key into the store (admin configuration page)."""
if name not in TOOL_KEY_NAMES:
raise ValueError(f"Clé non prise en charge: {name}")
value = (value or "").strip()
keys = _read_keys()
if value:
keys[name] = value
else:
keys.pop(name, None)
_write_keys(keys)
def delete_tool_key(name: str) -> bool:
"""Remove one key from the store; return True when it existed."""
if name not in TOOL_KEY_NAMES:
raise ValueError(f"Clé non prise en charge: {name}")
keys = _read_keys()
if name in keys:
del keys[name]
_write_keys(keys)
return True
return False
+181 -4
View File
@@ -8,6 +8,16 @@ Phase 1 of the documented web-toolset roadmap:
answering « je n'ai pas accès à internet ».
* ``fetch_url`` — retrieve a public web page and return readable text.
Phase 2 (#92) additions:
* keyed providers — Tavily, Brave Search, SerpAPI and Exa are used first when
their API key is configured (env, injected by Infisical in production);
* SQLite cache — search/fetch results are cached with a TTL
(:mod:`backend.tools.webcache`);
* retry with backoff — transient network errors get one extra attempt;
* dynamic rendering — ``fetch_url(render=True)`` uses an isolated Playwright
worker (optional dependency, graceful degradation).
All are READ-risk tools (no confirmation), rate-limited through the shared
registry, SSRF-guarded (scheme + private-address rejection), and size-capped.
@@ -16,6 +26,14 @@ Configuration (environment):
* ``OBSIGATE_WEB_TIMEOUT`` — seconds, default 10
* ``OBSIGATE_WEB_FALLBACK`` — ``0``/``false`` disables the keyless HTML
fallbacks (SearXNG only), default enabled
* ``OBSIGATE_TAVILY_API_KEY`` / ``OBSIGATE_BRAVE_API_KEY`` /
``OBSIGATE_SERPAPI_API_KEY`` / ``OBSIGATE_EXA_API_KEY`` — optional keyed
providers, tried before SearXNG when set
* ``OBSIGATE_WEB_PROVIDERS`` — optional comma-separated provider order
(e.g. ``brave,searxng``); keyed providers without a key are skipped
* ``OBSIGATE_WEB_RETRY`` — extra attempts for transient network errors
(default 1)
* ``OBSIGATE_WEB_CACHE_TTL`` — cache TTL seconds, ``0`` disables (default 900)
"""
from __future__ import annotations
@@ -28,15 +46,18 @@ import logging
import os
import re
import socket
import time
from collections.abc import Callable
from typing import Any
from urllib.parse import parse_qs, urlparse
import httpx
from backend.tools import webcache
from backend.tools.context import ToolError, ToolRisk, ToolScope
from backend.tools.registry import tool
from backend.tools.schemas import FetchUrlInput, WebSearchInput
from backend.tools.secrets import get_tool_key
logger = logging.getLogger("obsigate.tools.web")
@@ -48,6 +69,7 @@ WEB_FALLBACK_ENABLED = os.environ.get("OBSIGATE_WEB_FALLBACK", "1").strip().lowe
"no",
"off",
}
WEB_RETRY_ATTEMPTS = int(os.environ.get("OBSIGATE_WEB_RETRY", "1"))
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
# Search engines reject non-browser agents on their public HTML endpoints.
BROWSER_UA = (
@@ -163,6 +185,115 @@ def _result(
}
def _with_retry(call: Callable[[], Any]) -> Any:
"""Run *call* with one extra attempt on transient network errors.
House-made backoff (the roadmap's « tenacity ou boucle maison »): DNS
blips and rate-limit hiccups are the common failure mode, and a single
retry keeps the fallback chain from being consumed too early.
"""
for attempt in range(1 + max(0, WEB_RETRY_ATTEMPTS)):
try:
return call()
except httpx.TransportError:
if attempt >= max(0, WEB_RETRY_ATTEMPTS):
raise
time.sleep(0.2 * (attempt + 1))
raise RuntimeError("unreachable") # pragma: no cover
def _env_key(name: str) -> str:
"""Read an API key: configuration-page store first, then environment."""
return get_tool_key(name)
def _search_tavily(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Tavily Search API (agent-oriented results, key required)."""
resp = httpx.post(
"https://api.tavily.com/search",
json={
"api_key": _env_key("OBSIGATE_TAVILY_API_KEY"),
"query": query,
"max_results": params.max_results,
"search_depth": "basic",
"include_answer": False,
},
headers={"User-Agent": USER_AGENT},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", item.get("content") or "")
for item in (data.get("results") or [])
], []
def _search_brave(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Brave Search API (key required)."""
resp = httpx.get(
"https://api.search.brave.com/res/v1/web/search",
params={"q": query, "count": params.max_results, "safesearch": "moderate"},
headers={
"X-Subscription-Id": _env_key("OBSIGATE_BRAVE_API_KEY"),
"Accept": "application/json",
"User-Agent": USER_AGENT,
},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", item.get("description") or "")
for item in ((data.get("web") or {}).get("results") or [])
], []
def _search_serpapi(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""SerpAPI (Google SERP, key required)."""
resp = httpx.get(
"https://serpapi.com/search",
params={"q": query, "api_key": _env_key("OBSIGATE_SERPAPI_API_KEY"),
"num": params.max_results},
headers={"User-Agent": USER_AGENT},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("link") or "", item.get("snippet") or "")
for item in (data.get("organic_results") or [])
], []
def _search_exa(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Exa neural search (key required)."""
resp = httpx.post(
"https://api.exa.ai/search",
json={"query": query, "numResults": params.max_results},
headers={
"x-api-key": _env_key("OBSIGATE_EXA_API_KEY"),
"User-Agent": USER_AGENT,
},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", (item.get("text") or "")[:600])
for item in (data.get("results") or [])
], []
# Keyed providers: name -> (implementation, API key env var)
_KEYED_PROVIDERS: dict[str, tuple[_Provider, str]] = {
"tavily": (_search_tavily, "OBSIGATE_TAVILY_API_KEY"),
"brave": (_search_brave, "OBSIGATE_BRAVE_API_KEY"),
"serpapi": (_search_serpapi, "OBSIGATE_SERPAPI_API_KEY"),
"exa": (_search_exa, "OBSIGATE_EXA_API_KEY"),
}
def _search_searxng(
query: str, params: WebSearchInput
) -> tuple[list[dict[str, Any]], list[str]]:
@@ -288,8 +419,22 @@ _Provider = Callable[[str, WebSearchInput], "tuple[list[dict[str, Any]], list[st
def _provider_chain() -> list[tuple[str, _Provider]]:
"""Ordered providers: self-hosted meta-search first, then keyless fallbacks."""
chain: list[tuple[str, _Provider]] = [("searxng", _search_searxng)]
"""Ordered providers: keyed APIs first, then self-hosted, then keyless.
``OBSIGATE_WEB_PROVIDERS`` (comma-separated) overrides the default order;
unknown names are ignored and keyed providers without their key are skipped.
"""
chain: list[tuple[str, _Provider]] = []
configured = [
name.strip().lower()
for name in os.environ.get("OBSIGATE_WEB_PROVIDERS", "").split(",")
if name.strip()
]
for name in configured or list(_KEYED_PROVIDERS):
entry = _KEYED_PROVIDERS.get(name)
if entry and _env_key(entry[1]):
chain.append((name, entry[0]))
chain.append(("searxng", _search_searxng))
if WEB_FALLBACK_ENABLED:
chain.append(("duckduckgo", _search_duckduckgo))
chain.append(("bing", _search_bing))
@@ -313,6 +458,17 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
if not query:
raise ToolError("Requête vide", code="invalid_arguments")
key = webcache.cache_key("search", {
"q": query,
"max_results": params.max_results,
"category": params.category,
"language": params.language,
"page": params.page,
})
cached = webcache.cache_get(key)
if cached is not None:
return {**cached, "cached": True}
attempts: list[str] = []
unresponsive: list[str] = []
reachable = False
@@ -320,8 +476,12 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
for name, provider in _provider_chain():
attempts.append(name)
def _attempt(p: _Provider = provider) -> tuple[list[dict[str, Any]], list[str]]:
return p(query, params)
try:
results, engines = provider(query, params)
results, engines = _with_retry(_attempt)
except (httpx.HTTPError, ValueError, AttributeError) as e:
logger.warning("web_search provider %s failed: %s", name, e)
last_error = e
@@ -338,6 +498,7 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
}
if unresponsive:
payload["unresponsive_engines"] = unresponsive[:8]
webcache.cache_set(key, payload)
return payload
if not reachable:
@@ -378,6 +539,20 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
"""Retrieve one page, guard against SSRF, and extract its text."""
url = _assert_public_http_url(params.url.strip())
key = webcache.cache_key("fetch", {"url": url, "render": params.render})
cached = webcache.cache_get(key)
if cached is not None:
return {**cached, "cached": True}
if params.render:
# Dynamic pages (SPA/React): delegated to the isolated Playwright
# worker; the browser dependency stays optional (graceful error).
from backend.tools.webrender import render_page
payload = render_page(url)
webcache.cache_set(key, payload)
return payload
try:
# Follow redirects manually so every hop is re-checked against the
# private-address SSRF guard (a public page can redirect to 127.0.0.1).
@@ -414,10 +589,12 @@ def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
title_match = re.search(r"<title[^>]*>(.*?)</title>", raw, re.IGNORECASE | re.DOTALL)
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
return {
payload = {
"url": str(resp.url),
"status": resp.status_code,
"title": title,
"text": text,
"truncated": len(raw) > MAX_TEXT_CHARS,
}
webcache.cache_set(key, payload)
return payload
+138
View File
@@ -0,0 +1,138 @@
"""SQLite cache for web tool results (search results, fetched pages).
Phase 2 of the web-toolset roadmap (« Transverse »): repeated web searches and
page fetches (common in agent loops, where the model re-reads a source) must
not hammer the providers. Results are cached in a dedicated SQLite table with
a TTL; the cache is best-effort — any error silently disables it so a broken
database file never takes the assistant down.
Configuration (environment):
* ``OBSIGATE_DATA_DIR`` — base data directory (default ``data``)
* ``OBSIGATE_WEB_CACHE_PATH`` — explicit cache file override
* ``OBSIGATE_WEB_CACHE_TTL`` — seconds, ``0`` disables the cache (default 900)
"""
from __future__ import annotations
import hashlib
import json
import logging
import os
import sqlite3
import threading
import time
from pathlib import Path
from typing import Any
logger = logging.getLogger("obsigate.tools.webcache")
DEFAULT_TTL_SECONDS = 900
_schema_ready = False
_write_lock = threading.Lock()
def ttl_seconds() -> float:
"""Configured TTL in seconds (``0`` = cache disabled)."""
return float(os.environ.get("OBSIGATE_WEB_CACHE_TTL", str(DEFAULT_TTL_SECONDS)))
def _cache_path() -> Path:
override = os.environ.get("OBSIGATE_WEB_CACHE_PATH", "").strip()
if override:
return Path(override)
return Path(os.environ.get("OBSIGATE_DATA_DIR", "data")) / "web_cache.sqlite3"
def _connect() -> sqlite3.Connection:
"""Open (and lazily create) the cache database."""
global _schema_ready
path = _cache_path()
path.parent.mkdir(parents=True, exist_ok=True)
conn = sqlite3.connect(path, timeout=5, check_same_thread=False)
if not _schema_ready:
conn.execute(
"CREATE TABLE IF NOT EXISTS web_cache ("
"key TEXT PRIMARY KEY, value TEXT NOT NULL, created REAL NOT NULL)"
)
conn.commit()
_schema_ready = True
return conn
def cache_key(prefix: str, payload: dict[str, Any]) -> str:
"""Deterministic cache key from a prefix and the normalized arguments."""
raw = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str)
digest = hashlib.sha256(raw.encode("utf-8")).hexdigest()[:32]
return f"{prefix}:{digest}"
def cache_get(key: str) -> Any | None:
"""Return the cached payload for *key*, or ``None`` (miss/expiry/disabled)."""
if ttl_seconds() <= 0:
return None
try:
conn = _connect()
row = conn.execute(
"SELECT value, created FROM web_cache WHERE key = ?", (key,)
).fetchone()
conn.close()
except sqlite3.Error as e:
logger.warning("web cache read failed (%s): %s", key, e)
return None
if row is None:
return None
value, created = row
if time.time() - float(created) > ttl_seconds():
return None
try:
return json.loads(value)
except (ValueError, TypeError):
return None
def cache_set(key: str, value: Any) -> None:
"""Store *value* under *key* (best effort, never raises)."""
if ttl_seconds() <= 0:
return
try:
with _write_lock:
conn = _connect()
conn.execute(
"INSERT INTO web_cache (key, value, created) VALUES (?, ?, ?) "
"ON CONFLICT(key) DO UPDATE SET value = excluded.value, created = excluded.created",
(key, json.dumps(value, ensure_ascii=False, default=str), time.time()),
)
conn.commit()
conn.close()
except sqlite3.Error as e:
logger.warning("web cache write failed (%s): %s", key, e)
def purge_expired() -> int:
"""Delete expired rows; return the number of removed entries (maintenance)."""
try:
conn = _connect()
cursor = conn.execute(
"DELETE FROM web_cache WHERE created < ?", (time.time() - ttl_seconds(),)
)
conn.commit()
deleted = cursor.rowcount
conn.close()
return int(deleted)
except sqlite3.Error as e:
logger.warning("web cache purge failed: %s", e)
return 0
def clear_cache() -> int:
"""Drop every cached entry (tests / admin); returns the number of rows."""
try:
conn = _connect()
cursor = conn.execute("DELETE FROM web_cache")
conn.commit()
deleted = cursor.rowcount
conn.close()
return int(deleted)
except sqlite3.Error as e:
logger.warning("web cache clear failed: %s", e)
return 0
+100
View File
@@ -0,0 +1,100 @@
"""Dynamic page rendering (Playwright) — ``fetch_url(render=True)``.
Static pages are fetched with httpx inside :mod:`backend.tools.web`. Dynamic
pages (SPA/React, JS-loaded content) need a real browser engine; this module
runs one Playwright call inside a dedicated worker thread so browser
crashes/timeouts never take over the tool layer, and the heavyweight
dependency stays optional:
* not installed → ``ToolError(code="playwright_unavailable")`` with a clear
message (the assistant explains the limitation instead of hanging);
* installed → ``pip install playwright && playwright install chromium``.
The SSRF guard (scheme + private-address rejection) is applied before the
browser navigates. Note: unlike the httpx path, internal redirects performed
by the browser engine are not re-checked hop by hop.
"""
from __future__ import annotations
import html as html_lib
import logging
import re
from concurrent.futures import ThreadPoolExecutor
from typing import Any
from backend.tools.context import ToolError
from backend.tools.web import (
MAX_TEXT_CHARS,
USER_AGENT,
_assert_public_http_url,
_html_to_text,
)
logger = logging.getLogger("obsigate.tools.webrender")
# One worker: browser automation is serialized on purpose (one Chromium at a
# time keeps memory predictable on small hosts).
_executor = ThreadPoolExecutor(max_workers=1, thread_name_prefix="obsigate-playwright")
GOTO_TIMEOUT_MS = 20_000
def _playwright_available() -> bool:
try:
import playwright # noqa: F401
except ImportError:
return False
return True
def _render_in_worker(url: str) -> dict[str, Any]:
"""Synchronous Playwright render — runs in the dedicated worker thread."""
from playwright.sync_api import sync_playwright
status = 0
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
page = browser.new_page(user_agent=USER_AGENT)
response = page.goto(url, wait_until="networkidle", timeout=GOTO_TIMEOUT_MS)
if response is not None:
status = response.status
raw = page.content()
title = html_lib.unescape(page.title() or "").strip()
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
finally:
browser.close()
title = re.sub(r"\s+", " ", title)[:300]
return {
"url": url,
"status": status,
"title": title,
"text": text,
"rendered": True,
"truncated": len(raw) > MAX_TEXT_CHARS,
}
def render_page(url: str) -> dict[str, Any]:
"""Render *url* (JavaScript included) and return readable text.
Raises:
ToolError: ``playwright_unavailable`` when the optional dependency is
missing, ``render_unavailable`` when the render itself failed.
"""
_assert_public_http_url(url)
if not _playwright_available():
raise ToolError(
"Rendu dynamique indisponible : Playwright n'est pas installé "
"(pip install playwright && playwright install chromium).",
code="playwright_unavailable",
)
try:
return _executor.submit(_render_in_worker, url).result(timeout=GOTO_TIMEOUT_MS / 1000 + 40)
except ToolError:
raise
except Exception as e:
logger.warning("render_page failed for %s: %s", url, e)
raise ToolError(
"Le rendu dynamique de la page a échoué.", code="render_unavailable"
) from e
+1 -1
View File
@@ -2626,7 +2626,7 @@ dependencies = [
[[package]]
name = "obsigate-desktop"
version = "2.9.1"
version = "2.11.2"
dependencies = [
"chrono",
"env_logger",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "obsigate-desktop"
version = "2.9.1"
version = "2.11.2"
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
authors = ["Bruno Charest"]
edition = "2021"
+1 -1
View File
@@ -1,7 +1,7 @@
{
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
"productName": "ObsiGate",
"version": "2.9.1",
"version": "2.11.2",
"identifier": "com.obsigate.desktop",
"build": {
"frontendDist": "../frontend",
+3
View File
@@ -168,6 +168,7 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-056* | [🟡 IMPORTANT] Éditeur Forge en plein écran : l'Assistant IA s'ouvre en arrière-plan et reste invisible | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs` | Forge : passer en plein écran puis cliquer le bouton « Assistant IA » (ou `Ctrl+J`) — le panneau s'ouvre dans le document parent, masqué par l'iframe plein écran | Sortie du plein écran **avant** d'ouvrir le panneau, des deux côtés : côté iframe (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et** côté parent (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut être détenu par le document parent et non par l'iframe (dans ce cas `document.fullscreenElement` est nul dans l'iframe et sa sortie échoue). Tests : `forge-completion.test.mjs` (+1), `editor-inline.test.mjs` (+1) | Le panneau assistant est monté dans `document.body` du parent : l'API Fullscreen ne rend que l'élément plein écran et ses descendants, donc il ne peut pas s'afficher au-dessus de l'iframe Forge en plein écran. La sortie côté iframe seule ne suffisait pas quand le parent détient le plein écran |
| *BUG-057* | [🟡 IMPORTANT] Assistant IA : le bouton « Ajouter » est inopérant dans l'éditeur Forge (fonctionne seulement dans « Editer ») | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js`, `frontend/editor-poc.html` | Ouvrir un document dans Forge, demander une réponse à l'assistant puis cliquer « Ajouter » | `_insertIntoEditor()` cible Forge (`#forge-iframe`) : `postMessage({ type: 'parent-insert', text })` ; `editor-poc.html` insère au curseur (`insertAtCursor`) et marque le tampon modifié. Repli textarea inclus. Tests : `tests/frontend/ai.test.mjs` (+3), `tests/frontend/editor-inline.test.mjs` (+1) | `state.editorView` (CodeMirror) est nul en Forge : le clic affichait « Aucun document ouvert dans l'éditeur » |
| *BUG-058* | [🔵 MINEUR] Éditeur « Editer » : la barre de numérotation de ligne ne suit pas la couleur du thème (gutter clair `#f5f5f5` en thème sombre) | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css` | Ouvrir un document → Editer en thème sombre : la colonne des numéros de ligne reste gris clair alors que le fond de l'éditeur est sombre | Thème du gutter CodeMirror via les variables CSS (`color-mix(var(--text-primary) …)` pour le fond, `--text-secondary` pour les numéros, `--border` pour la séparation, `--text-primary` pour la ligne active) au lieu des valeurs codées en dur de CodeMirror ; test de non-régression dans `tests/frontend/editor-inline.test.mjs`. Vérifié Playwright (instance de test) : sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de` | CodeMirror applique `background:#f5f5f5` par défaut, indépendamment du thème ObsiGate ; en mode sombre le fond de l'éditeur suit `--bg-secondary` mais pas le gutter |
| *BUG-060* | [🟡 IMPORTANT] Viewer PDF : l'affichage des pages ne fonctionne pas — seule la barre d'outils « PDF — N pages » s'affiche, le contenu reste vide | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau) | Cliquer un fichier `.pdf` dans l'arborescence | `frontend/js/viewer.js` : le rendu PDF passe de `<embed type="application/pdf">` à `<iframe>` (autorisée par `frame-src 'self'`, le stream étant same-origin). Tests : `tests/frontend/pdf-viewer.test.mjs` (+6) et `tests/e2e/pdf-viewer.spec.js` (fixture `test_vault/sample-pdf.pdf`) | Cause : la CSP durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>` → le lecteur PDF natif était bloqué (barre d'outils rendue, corps vide). Le test E2E échoue bien avec l'ancien `<embed>`. `object-src 'none'` conservé (le correctif ne désarme pas la CSP) |
| | | | | | | | | | | |
### TODOs techniques (améliorations / nouvelles tâches)
@@ -229,6 +230,8 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| 2026-09-17 | BUG-056 | Correction | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-056** : en plein écran Forge, l'Assistant IA s'ouvrait en arrière-plan. La sortie du plein écran est désormais faite **côté iframe** (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et côté parent** (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut appartenir au document parent (l'iframe voit alors `fullscreenElement` nul et sa sortie échoue — c'était le cas non couvert par le premier correctif). Vérifié : `forge-completion.test.mjs` 32/32 (+1), `editor-inline.test.mjs` 42/42 (+1), 14 suites frontend vertes, validate-imports 38 modules. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-057, #102 | Correction + feature | `frontend/js/bookslm.js`, `frontend/editor-poc.html`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/frontend/ai.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `docs/archive/COMPLETED_v1-v2.md`, `docs/ROADMAP.md`, `CHANGELOG.md` | **BUG-057** : le bouton « Ajouter » de l'assistant ne ciblait que `state.editorView` (CodeMirror) ; en Forge il affichait « Aucun document ouvert dans l'éditeur ». `_insertIntoEditor()` gère désormais les trois surfaces : CodeMirror, l'iframe Forge (`postMessage({ type: 'parent-insert', text })` → `insertAtCursor` dans `editor-poc.html`) et le textarea de repli. **#102** : chaque bloc de code d'une réponse reçoit un bouton « Ajouter la section » (`.bookslm-code-insert`, révélé au survol) qui insère le contenu du bloc sans les délimiteurs ` ``` `. Vérifié : `ai.test.mjs` 91/91 (+3), `editor-inline.test.mjs` 43/43 (+1), `forge-completion.test.mjs` 32/32, unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-058 | Correction | `frontend/style.css`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-058** : la barre de numérotation de ligne de l'éditeur « Editer » ne suivait pas le thème — CodeMirror peint `.cm-gutters` avec des valeurs claires codées en dur (`#f5f5f5`, bordure `#ddd`), visibles en thème sombre. Correctif : le gutter dérive des variables CSS ObsiGate (`background: color-mix(in srgb, var(--text-primary) 5%, transparent)`, `color: var(--text-secondary)`, `border-right: 1px solid var(--border)`, ligne active `color-mix(… 10% …)` / `--text-primary`), donc il suit les 15 thèmes et les 4 modes. Vérifié : `editor-inline.test.mjs` 44/44 (+1), unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0, et Playwright sur l'instance de test (route `style.css` remplacée par le fichier local) — sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de`, plus de `rgb(245,245,245)`. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-059 | Correction | `frontend/js/bookslm.js`, `tests/frontend/ai.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-059** : dans une conversation ouverte (post ancré en haut), **tout clic** dans la fenêtre de messages — lien de fichier, étapes, sélection de texte — faisait sauter toute la conversation au bas de la fenêtre. Cause : le gestionnaire `mousedown` de dépintage (prévu pour la molette/tactile/poignée de scroll) se déclenchait aussi sur un simple clic, et le retrait du padding d'ancre (`paddingBottom`) bornait le `scrollTop` à la nouvelle hauteur max → saut au bas. Correctif : helper pur `isScrollbarPress(target, clientX, container)` — un appui ne dépine que s'il vise la **poignée de scroll** (cible = conteneur + zone de gouttière droite) ; molette et tactile conservent leur comportement. Vérifié : `ai.test.mjs` 92/92 (+1), unit 9/9, validate-imports 38 modules, pytest / ruff / mypy inchangés côté backend. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-060 | Correction | `frontend/js/viewer.js`, `.gitea/workflows/ci.yml`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau), `test_vault/sample-pdf.pdf` (nouveau), `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-060** : l'ouverture d'un PDF n'affichait aucune page (barre d'outils « PDF — N pages » présente, corps vide). Cause : la CSP durcie en BUG-034 pose `object-src 'none'` — directive qui gouverne `<embed>`/`<object>` — alors que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était bloqué. Correctif : rendu dans une `<iframe>` (autorisée par `frame-src 'self'`, le stream `/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` conservé. Tests : `pdf-viewer.test.mjs` 6/6 (statique : pas d'`<embed>`, CSP `frame-src 'self'`, iframe pleine hauteur), `pdf-viewer.spec.js` (E2E : iframe + stream `application/pdf` 200/206 + zéro violation CSP ; échoue bien avec l'ancien `<embed>`). Vérifié : pytest 1184 passed / 6 skipped, frontend 14 suites JSDOM vertes, validate-imports 38 modules, ruff/mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
---
+5 -21
View File
@@ -1,6 +1,6 @@
# ObsiGate — Roadmap
> **Version :** 2.9.1 | **Dernière mise à jour :** 2026-09-17
> **Version :** 2.11.2 | **Dernière mise à jour :** 2026-09-17
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
> vers les fonctionnalités livrées.
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
@@ -78,23 +78,6 @@
- [x] Personnalisation (clé à molette) : ajouter / supprimer / réordonner les commandes
- [x] i18n FR/EN + tests frontend (helpers purs) + E2E mobile
### 92. Assistant IA — Écosystème d'outils (phase 2 : web étendu, sources connectées, documents)
- **Effort :** 3-5 jours | **Impact :** 🟠 | **Zone :** backend (`backend/tools/`)
- **Dépend de :** #91 (registre + section « steps » + `web_search`/`fetch_url` livrés)
- **Description :** étendre le catalogue d'outils de l'assistant au-delà du vault, en
suivant la feuille de route technique détaillée :
[features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) (frameworks évalués,
bibliothèques par catégorie, transverse retry/cache/secrets/async).
- **Sous-tâches :**
- [x] `web_search` : chaîne de repli sans clé (SearXNG → DuckDuckGo → Bing, `OBSIGATE_WEB_FALLBACK`) — BUG-051
- [ ] `web_search` : fournisseurs optionnels à clé (Tavily, Brave, SerpAPI, Exa)
- [ ] `fetch_url` : pages dynamiques via Playwright (worker isolé) ; crawl multi-pages Scrapy en tâche de fond
- [ ] Sources connectées : Gitea/GitHub (priorité haute) puis Google Drive / OneDrive (OAuth2 `authlib`)
- [ ] Production de documents : conversion, tableurs, PDF/Word (outils WRITE + confirmation)
- [ ] Transverse : `tenacity` (backoff), cache SQLite des résultats web avec TTL, secrets via Infisical
- [ ] Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests (httpx mocké)
---
## ⚪ Backlog — Sécurité, architecture & performance (P0/P1)
@@ -201,6 +184,8 @@
| 101 | Forge — Assistant IA partagé (bouton AI Panel = assistant, fournisseur/modèle configuré, autocomplétion) + plein écran Forge/Editer | 2.8.0 | [features/forge-assistant.md](./features/forge-assistant.md) |
| BUG-057 | Assistant IA — bouton « Ajouter » fonctionnel dans l'éditeur Forge (en plus d'« Editer ») | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
| 102 | Assistant IA — bouton « Ajouter la section » par bloc de code (insertion du bloc seul) | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
| 92 | Assistant IA — Écosystème d'outils phase 2 (recherche à clé, cache/retry, Playwright, crawl, Gitea/GitHub, documents XLSX/DOCX/CSV/PDF) | 2.10.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
| 103 | Configuration — clés utilisateur des sources connectées & recherche à clé (page Configurations, `data/api_keys.json`, priorité sur l'env) | 2.11.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
---
@@ -208,12 +193,11 @@
| Priorité | Items | Effort total estimé |
|---|---|---|
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100 | ~114 jours réalisés |
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100, #102, #92 | ~114 jours réalisés |
| 🔵 P2 restant | #77 Desktop : signature de code (non retenue), 6 tests E2E **manuels** ([protocole](./DESKTOP_E2E_CHECKLIST.md)) | ~0,5-1 jour |
| ⚪ P4 restant | #73 Sync (6-8j) | 6-8 jours |
| ⚪ P2 restant | #92 Assistant IA — écosystème d'outils phase 2 (web étendu, sources connectées, documents) | 3-5 jours |
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (issues BUG-035 → BUG-040) | ~15-23 jours |
| **Total restant** | **10 items + finitions** | **~30-47 jours** |
| **Total restant** | **7 items + finitions** | **~27-42 jours** |
---
+12 -1
View File
@@ -1,6 +1,6 @@
# #92 — Assistant IA — Écosystème d'outils : feuille de route technique
> **Statut :** ⚪ Backlog (phase 1 livrée dans #91)
> **Statut :** ✅ livré (phase 2, version 2.10.0) — phase 1 livrée dans #91
> **Effort estimé :** 3-5 jours pour la phase 2 | **Impact :** 🟠
> **Références :** [Roadmap](../ROADMAP.md) · [Outils & MCP #79](./ai-tools-mcp.md) ·
> [Fenêtre de discussion #91](./ai-assistant-conversation-ux.md) · [Changelog](../../CHANGELOG.md)
@@ -36,6 +36,17 @@ réécriture de la boucle n'est nécessaire.
## 3. Phase 2 — catégories à implémenter
> **Livré (2.10.0, #92).** Récapitulatif des décisions finales :
| Catégorie | Décision livrée |
|---|---|
| Recherche web étendue | Tavily, Brave, SerpAPI, Exa à clé (`OBSIGATE_*_API_KEY`), essayés avant SearXNG ; ordre via `OBSIGATE_WEB_PROVIDERS` |
| Lecture de pages | `fetch_url(render=True)` → worker Playwright isolé (`backend/tools/webrender.py`), dépendance optionnelle + erreur explicite |
| Crawl multi-pages | `crawl_site` (WRITE + confirmation) : BFS httpx borné (≤ 20 pages, même hôte, SSRF sur chaque URL) → condensé Markdown dans le vault. Scrapy écarté (dépendance lourde inutile à cette échelle) |
| Sources connectées | Gitea + GitHub (`git_list_repos`, `git_search_issues`, `git_get_file`) via env/Infisical ; drives cloud (Drive/OneDrive) orientés serveur MCP externe (#79). **#103 (2.11.0)** : les clés (URL Gitea, tokens Gitea/GitHub, clés Tavily/Brave/SerpAPI/Exa) se saisissent aussi dans la page Configurations — `backend/tools/secrets.py`, valeur stockée prioritaire sur l'env |
| Production de documents | `create_xlsx`, `create_docx`, `create_csv`, `create_pdf` — WRITE + confirmation, écrit via `save_raw_file(allow_docs=True)` (path safety + backup) |
| Transverse | Cache SQLite (`webcache.py`, TTL `OBSIGATE_WEB_CACHE_TTL`), retry backoff maison (`OBSIGATE_WEB_RETRY`), secrets par env (Infisical-compatible) |
### 3.1 Recherche web étendue (`web_search`)
- **Fallback sans clé — ✅ livré (BUG-051)** : chaîne de fournisseurs dans
`backend/tools/web.py` — SearXNG auto-hébergé (`OBSIGATE_SEARXNG_URL`) puis, si
+1
View File
@@ -8,6 +8,7 @@
- **Implémentation réelle (vérifiée 2026-09-07) :**
- **Bugs corrigés (2026-09) :** `api_pdf_stream` crashait en 500 (`NameError: current_user` jamais injecté) ; l'indexation incrémentale du watcher faisait `read_text()` sur les PDFs (garbage) ; Range/206 et `pdf/info` absents malgré le texte ci-dessous.
- **BUG-060 (2026-09-17) :** l'affichage inline ne fonctionnait plus — la CSP durcie en BUG-034 (`object-src 'none'`) bloquait l'`<embed>` du viewer (barre d'outils rendue, corps vide). Le rendu passe par une `<iframe>` (autorisée par `frame-src 'self'`), conforme à E1. Tests : `tests/frontend/pdf-viewer.test.mjs` + `tests/e2e/pdf-viewer.spec.js`.
- `GET /api/file/{vault}/pdf/info` — métadonnées seules sans transférer le document (C3)
- Stream avec `Accept-Ranges` + 206 Partial Content (single range, suffix-range, 416) (C2)
- `OBSIGATE_PDF_MAX_SIZE_MB` (50) + `OBSIGATE_PDF_EXTRACT_TIMEOUT` (30s via thread-pool) (B4/G3)
+128
View File
@@ -1543,6 +1543,7 @@
<li><a href="#cfg-hidden-files" class="help-nav-link" data-i18n="config.section_hidden"></a></li>
<li><a href="#cfg-diags" class="help-nav-link" data-i18n="settings.diagnostics"></a></li>
<li><a href="#cfg-ai" class="help-nav-link" data-i18n="settings.ai"></a></li>
<li><a href="#cfg-sources" class="help-nav-link" data-i18n="config.section_sources">Sources connectées</a></li>
<li><a href="#cfg-themes" class="help-nav-link" data-i18n="settings.themes"></a></li>
<li><a href="#cfg-profile" class="help-nav-link" data-i18n="settings.profile"></a></li>
<li><a href="#cfg-security" class="help-nav-link" data-i18n="settings.security"></a></li>
@@ -2260,6 +2261,133 @@
</div>
</section>
<!-- Connected sources & keyed web search (#103) -->
<section
class="config-section help-section"
id="cfg-sources"
>
<h2 data-i18n="config.section_sources"
>Sources connectées & recherche web</h2
>
<p
class="config-description"
data-i18n="config.sources_desc"
>
Clés utilisées par les outils de l'Assistant IA
(recherche web à clé, Gitea, GitHub). Elles sont
stockées localement et priment sur les variables
d'environnement.
</p>
<div class="config-row">
<label
class="config-label"
for="cfg-tavily-key"
>Tavily API Key</label
>
<input
type="password"
id="cfg-tavily-key"
class="config-input"
placeholder="tvly-..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-brave-key"
>Brave Search API Key</label
>
<input
type="password"
id="cfg-brave-key"
class="config-input"
placeholder="BSA..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-serpapi-key"
>SerpAPI Key</label
>
<input
type="password"
id="cfg-serpapi-key"
class="config-input"
placeholder="..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-exa-key"
>Exa API Key</label
>
<input
type="password"
id="cfg-exa-key"
class="config-input"
placeholder="..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-gitea-url"
data-i18n="config.gitea_url"
>URL Gitea</label
>
<input
type="text"
id="cfg-gitea-url"
class="config-input"
placeholder="https://git.example.net"
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-gitea-token"
>Gitea Token</label
>
<input
type="password"
id="cfg-gitea-token"
class="config-input"
placeholder="token personnel..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-github-token"
>GitHub Token</label
>
<input
type="password"
id="cfg-github-token"
class="config-input"
placeholder="ghp_..."
autocomplete="off"
/>
</div>
<div
class="config-actions-row"
style="margin-top: 16px"
>
<button class="config-btn-save"
id="cfg-save-tool-keys" data-i18n="help.shortcut_save">
Sauvegarder
</button>
</div>
</section>
<!-- Themes -->
<section
class="config-section help-section"
+25 -1
View File
@@ -43,6 +43,25 @@ const PANEL_MIN_WIDTH = 320;
const PANEL_MAX_WIDTH = 1000;
const PANEL_WIDTH_KEY = 'obsigate-bookslm-width';
/**
* BUG-059 — Is this pointer press a scrollbar drag (and only that)?
*
* Only genuine scroll gestures may release the top-pinning of the latest
* question (wheel / touchmove / scrollbar drag). A scrollbar drag reports the
* scrollable container itself as the event target and lands inside the
* vertical scrollbar gutter. A plain click anywhere in the content — a link,
* a file path, the steps toggle, text selection — must NOT unpin: clearing
* the anchor padding would clamp the scroll position and jump the whole
* thread to the bottom of the conversation window.
*/
export function isScrollbarPress(target, clientX, container) {
if (!container || target !== container) return false;
const rect = container.getBoundingClientRect();
if (!rect || !Number.isFinite(rect.right)) return false;
const GUTTER = 24; // conservative vertical-scrollbar width estimate
return clientX >= rect.right - GUTTER;
}
/**
* Accent- and case-insensitive normalization used to match `/` commands and
* `@` mentions: skill ids/labels and vault paths may contain accented
@@ -985,7 +1004,12 @@ class BooksLM {
};
messagesEl.addEventListener('wheel', unpin, { passive: true });
messagesEl.addEventListener('touchmove', unpin, { passive: true });
messagesEl.addEventListener('mousedown', unpin, { passive: true });
// BUG-059: a scrollbar drag only — a plain click on the content must
// not unpin (it would clear the anchor padding and jump the thread to
// the bottom of the window).
messagesEl.addEventListener('mousedown', (e) => {
if (isScrollbarPress(e.target, e.clientX, messagesEl)) unpin();
}, { passive: true });
}
// Close the session / command menus when clicking elsewhere in the panel.
+100
View File
@@ -715,6 +715,7 @@ function initConfigModal() {
await loadHiddenFilesSettings();
loadWebhooksUI();
loadSharesUI();
loadToolKeys();
safeCreateIcons();
});
@@ -765,6 +766,9 @@ function initConfigModal() {
if (saveAIKeysBtn) saveAIKeysBtn.addEventListener("click", saveAIKeys);
const testAIKeysBtn = document.getElementById("cfg-test-ai-keys");
if (testAIKeysBtn) testAIKeysBtn.addEventListener("click", testAIKeys);
// Tool & connected-source keys (#103)
const saveToolKeysBtn = document.getElementById("cfg-save-tool-keys");
if (saveToolKeysBtn) saveToolKeysBtn.addEventListener("click", saveToolKeys);
// Default provider/model selection
const aiDefaultProviderSel = document.getElementById("cfg-ai-default-provider");
if (aiDefaultProviderSel) {
@@ -1715,6 +1719,102 @@ async function testAIKeys() {
}
// ── Tool & connected-source keys (#103) ──
const TOOL_KEY_MAP = {
"cfg-tavily-key": "OBSIGATE_TAVILY_API_KEY",
"cfg-brave-key": "OBSIGATE_BRAVE_API_KEY",
"cfg-serpapi-key": "OBSIGATE_SERPAPI_API_KEY",
"cfg-exa-key": "OBSIGATE_EXA_API_KEY",
"cfg-gitea-url": "OBSIGATE_GITEA_URL",
"cfg-gitea-token": "OBSIGATE_GITEA_TOKEN",
"cfg-github-token": "OBSIGATE_GITHUB_TOKEN",
};
function _ensureToolKeyUI() {
for (const [inputId] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(inputId);
if (!input) continue;
const row = input.closest(".config-row");
if (!row || row.dataset.toolKeyEnhanced) continue;
row.dataset.toolKeyEnhanced = "1";
row.style.cssText += "display:flex;align-items:center;gap:8px;flex-wrap:wrap;";
const badge = document.createElement("span");
badge.id = inputId + "-badge";
badge.style.cssText = "font-size:11px;padding:2px 8px;border-radius:10px;white-space:nowrap;";
row.appendChild(badge);
const delBtn = document.createElement("button");
delBtn.type = "button";
delBtn.id = inputId + "-delete";
delBtn.className = "config-btn-secondary";
delBtn.style.cssText = "font-size:11px;padding:4px 10px;color:var(--danger,#e74c3c);border-color:var(--danger,#e74c3c);cursor:pointer;display:none;";
delBtn.textContent = "\u00d7 " + t("config.delete_key");
delBtn.addEventListener("click", () => deleteToolKey(inputId));
row.appendChild(delBtn);
}
}
function _setToolKeyBadge(inputId, hasKey) {
const badge = document.getElementById(inputId + "-badge");
const delBtn = document.getElementById(inputId + "-delete");
if (badge) {
if (hasKey) {
badge.textContent = "\u2713 " + t("config.key_set");
badge.style.background = "var(--success-bg, #27ae6022)";
badge.style.color = "var(--success, #27ae60)";
badge.style.border = "1px solid var(--success, #27ae60)";
} else {
badge.textContent = t("config.key_unset");
badge.style.background = "var(--muted-bg, #ffffff10)";
badge.style.color = "var(--text-muted, #888)";
badge.style.border = "1px solid var(--border, #444)";
}
}
if (delBtn) delBtn.style.display = hasKey ? "inline-block" : "none";
}
async function loadToolKeys() {
_ensureToolKeyUI();
try {
const data = await api("/api/config/tool-keys");
for (const [inputId, name] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(inputId);
const val = data[name] || "";
if (input && !input.value.trim()) input.placeholder = val || input.placeholder;
_setToolKeyBadge(inputId, !!val);
}
} catch(e) { /* non-admin: the section stays inert */ }
}
async function saveToolKeys() {
const keys = {};
for (const [id, name] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(id);
if (input && input.value.trim()) keys[name] = input.value.trim();
}
if (!Object.keys(keys).length) {
showToast(t("config.no_keys"), "warning");
return;
}
try {
await api("/api/config/tool-keys", { method: "POST", body: JSON.stringify(keys) });
showToast(t("config.api_keys_saved"), "success");
Object.keys(TOOL_KEY_MAP).forEach(id => { const el = document.getElementById(id); if (el) el.value = ""; });
loadToolKeys();
} catch(e) { showToast("Erreur: " + e.message, "error"); }
}
async function deleteToolKey(inputId) {
const name = TOOL_KEY_MAP[inputId];
if (!name) return;
if (!confirm(t("config.delete_key_confirm") + " " + name + " ?")) return;
try {
await api("/api/config/tool-keys/" + name, { method: "DELETE" });
showToast(t("config.key_deleted") + " " + name, "success");
loadToolKeys();
} catch(e) { showToast("Erreur: " + e.message, "error"); }
}
export {
initSidebarTabs,
initConfigModal,
+1 -1
View File
@@ -555,7 +555,7 @@ export function renderFile(data) {
</div>
<div class="pdf-body">
${tocHtml}
<embed src="${pdfUrl}" class="pdf-iframe" type="application/pdf" title="${escapeHtml(data.title)}"></embed>
<iframe src="${pdfUrl}" class="pdf-iframe" title="${escapeHtml(data.title)}"></iframe>
</div>
</div>`;
lucide.createIcons();
+16
View File
@@ -364,6 +364,14 @@
"config.ai_model": "Model",
"config.ai_openrouter_label": "OpenRouter API Key",
"config.api_keys_saved": "API keys saved",
"config.section_sources": "🔗 Connected sources & search",
"config.sources_desc": "Keys used by the AI assistant tools (keyed web search: Tavily, Brave, SerpAPI, Exa; connected sources: Gitea, GitHub). They are stored server-side and take precedence over environment variables.",
"config.gitea_url": "Gitea URL",
"config.key_set": "Configured",
"config.key_unset": "Not configured",
"config.delete_key": "Delete",
"config.delete_key_confirm": "Delete key",
"config.key_deleted": "Key deleted:",
"config.backups": "Backups",
"config.backups_desc": "Manage automatic file backups.",
"config.client_config": "Client config",
@@ -1822,6 +1830,14 @@
"ai.step.vaults": "Listed the vaults",
"ai.step.fetch_url": "Opened a web page: {value}",
"ai.step.web_search": "Searched the web: {value}",
"ai.step.crawl": "Crawled a site: {value}",
"ai.step.git_repos": "Listed repositories ({value})",
"ai.step.git_issues": "Searched issues: {value}",
"ai.step.git_file": "Read a repo file: {value}",
"ai.step.xlsx_create": "Spreadsheet proposed: {value}",
"ai.step.docx_create": "Word document proposed: {value}",
"ai.step.csv_create": "CSV file proposed: {value}",
"ai.step.pdf_create": "PDF document proposed: {value}",
"bookslm.copied": "Copied to clipboard",
"bookslm.error": "AI service error",
"bookslm.regenerate": "Regenerate",
+16
View File
@@ -364,6 +364,14 @@
"config.ai_model": "Modèle",
"config.ai_openrouter_label": "OpenRouter API Key",
"config.api_keys_saved": "Clés API sauvegardées",
"config.section_sources": "🔗 Sources connectées & recherche",
"config.sources_desc": "Clés utilisées par les outils de l'Assistant IA (recherche web à clé : Tavily, Brave, SerpAPI, Exa ; sources connectées : Gitea, GitHub). Elles sont stockées sur le serveur et priment sur les variables d'environnement.",
"config.gitea_url": "URL Gitea",
"config.key_set": "Configuré",
"config.key_unset": "Non configuré",
"config.delete_key": "Supprimer",
"config.delete_key_confirm": "Supprimer la clé",
"config.key_deleted": "Clé supprimée :",
"config.backups": "Sauvegardes",
"config.backups_desc": "Gérez les sauvegardes automatiques de vos fichiers.",
"config.client_config": "Configuration client",
@@ -1822,6 +1830,14 @@
"ai.step.vaults": "Liste des vaults consultée",
"ai.step.fetch_url": "Page web consultée : {value}",
"ai.step.web_search": "Recherche sur le web : {value}",
"ai.step.crawl": "Site exploré : {value}",
"ai.step.git_repos": "Dépôts listés ({value})",
"ai.step.git_issues": "Issues recherchées : {value}",
"ai.step.git_file": "Fichier de dépôt lu : {value}",
"ai.step.xlsx_create": "Tableur proposé : {value}",
"ai.step.docx_create": "Document Word proposé : {value}",
"ai.step.csv_create": "Fichier CSV proposé : {value}",
"ai.step.pdf_create": "Document PDF proposé : {value}",
"bookslm.copied": "Réponse copiée dans le presse-papiers",
"bookslm.error": "Erreur du service AI",
"bookslm.regenerate": "Régénérer",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "obsigate",
"version": "2.9.1",
"version": "2.11.2",
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
"main": "patch.js",
"directories": {
+106
View File
@@ -0,0 +1,106 @@
%PDF-1.3
%“Œ‹ž ReportLab Generated PDF document (opensource)
1 0 obj
<<
/F1 2 0 R
>>
endobj
2 0 obj
<<
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
>>
endobj
3 0 obj
<<
/Contents 9 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
4 0 obj
<<
/Contents 10 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
5 0 obj
<<
/Contents 11 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
6 0 obj
<<
/PageMode /UseNone /Pages 8 0 R /Type /Catalog
>>
endobj
7 0 obj
<<
/Author (anonymous) /CreationDate (D:20260917153218-04'00') /Creator (anonymous) /Keywords () /ModDate (D:20260917153218-04'00') /Producer (ReportLab PDF Library - \(opensource\))
/Subject (unspecified) /Title (untitled) /Trapped /False
>>
endobj
8 0 obj
<<
/Count 3 /Kids [ 3 0 R 4 0 R 5 0 R ] /Type /Pages
>>
endobj
9 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jt]&.8<uSu(.bn9(BF;Q*M*~>endstream
endobj
10 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jruos8<uSu(.bn9(BF;_*M3~>endstream
endobj
11 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?K!D1>8<uSu(.bn9(BF;m*M<~>endstream
endobj
xref
0 12
0000000000 65535 f
0000000061 00000 n
0000000092 00000 n
0000000199 00000 n
0000000392 00000 n
0000000586 00000 n
0000000780 00000 n
0000000848 00000 n
0000001109 00000 n
0000001180 00000 n
0000001405 00000 n
0000001631 00000 n
trailer
<<
/ID
[<464fc7cfdf793a5b6d29e3d0d043c5a7><464fc7cfdf793a5b6d29e3d0d043c5a7>]
% ReportLab generated PDF document -- digest (opensource)
/Info 7 0 R
/Root 6 0 R
/Size 12
>>
startxref
1857
%%EOF
+17
View File
@@ -22,6 +22,23 @@ def _reset_tool_ratelimit():
ratelimit.reset()
@pytest.fixture(autouse=True)
def _disable_web_cache():
"""Web cache off by default: tests stay hermetic (no cross-test hits).
tests/test_web_cache.py re-enables it explicitly with a tmp path.
"""
from backend.tools import webcache
saved_path = os.environ.get("OBSIGATE_WEB_CACHE_PATH")
os.environ["OBSIGATE_WEB_CACHE_TTL"] = "0"
yield
if saved_path is None:
os.environ.pop("OBSIGATE_WEB_CACHE_PATH", None)
else:
os.environ["OBSIGATE_WEB_CACHE_PATH"] = saved_path
@pytest.fixture(autouse=True)
def _clean_env():
"""Ensure no vault env vars leak between tests — but preserve test vault config."""
+92
View File
@@ -0,0 +1,92 @@
/**
* E2E tests for ObsiGate PDF inline viewer (BUG-060).
*
* Regression : la CSP posée par BUG-034 (`object-src 'none'`) bloque l'élément
* `<embed>` qui servait le PDF. Le viewer doit rendre le stream dans une
* `<iframe>` (autorisée par `frame-src 'self'`).
*
* Fixtures : `test_vault/sample-pdf.pdf` (3 pages, texte simple).
*
* Run (local):
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js --headed
*/
import { test, expect } from '@playwright/test';
const BASE = process.env.BASE_URL || 'http://localhost:2029';
const CREDS = {
username: process.env.OBSIGATE_USER || 'admin',
password: process.env.OBSIGATE_PASS || 'test123',
};
async function login(page) {
await page.goto(BASE);
const loginForm = page.locator('#login-screen');
await expect(loginForm).toBeVisible({ timeout: 5000 }).catch(() => {});
if (await loginForm.isVisible()) {
await page.fill('#login-username', CREDS.username);
await page.fill('#login-password', CREDS.password);
await page.click('#login-btn');
}
await page.waitForFunction(() => window.__OBSIGATE_BOOTED === true, { timeout: 20000 });
}
async function openFile(page, vault, filePath) {
const treeItem = page.locator(`.tree-item[data-vault="${vault}"][data-path="${filePath}"]`);
// Le vault peut être replié (auth activée) : l'étendre avant de chercher le fichier.
if (!(await treeItem.count())) {
await page.locator(`.tree-item.vault-item[data-vault="${vault}"]`).first().click();
await treeItem.waitFor({ state: 'attached', timeout: 8000 });
}
await treeItem.dblclick({ timeout: 5000 });
await page.waitForTimeout(500);
}
test.describe('PDF viewer — affichage inline (BUG-060)', () => {
test('ouvre un PDF dans une iframe (pas d\'<embed>) et le stream charge sans violation CSP', async ({ page }) => {
const cspViolations = [];
page.on('console', (msg) => {
const text = msg.text();
if (text.includes('Content Security Policy') && (text.includes('object-src') || text.includes('Refused'))) {
cspViolations.push(text);
}
});
await login(page);
// Le stream est demandé au moment où le viewer monte l'iframe : enregistrer
// l'écoute AVANT d'ouvrir le fichier (sinon la réponse est déjà passée).
const streamResponsePromise = page.waitForResponse(
(r) => r.url().includes('/pdf/stream') && (r.status() === 200 || r.status() === 206),
{ timeout: 15000 },
);
await openFile(page, 'TestVault', 'sample-pdf.pdf');
const iframe = page.locator('#content-area .pdf-iframe');
await expect(iframe).toBeVisible({ timeout: 15000 });
await expect(iframe).toHaveAttribute('src', /\/api\/file\/TestVault\/pdf\/stream\?path=/);
// La balise doit être une iframe (le <embed>/<object> serait bloqué par CSP)
const tagName = await iframe.evaluate((el) => el.tagName);
expect(tagName).toBe('IFRAME');
expect(await page.locator('#content-area embed, #content-area object').count()).toBe(0);
// La barre d'outils indique le nombre de pages du PDF
await expect(page.locator('.pdf-info')).toContainText('3 pages');
// Le stream est bien servi en application/pdf (200 ou 206 Range)
const streamResp = await streamResponsePromise;
expect(streamResp.headers()['content-type'] || '').toContain('application/pdf');
// Le cadre embarque réellement le document PDF (navigateur natif)
const pdfFrame = await iframe.contentFrame();
expect(pdfFrame).not.toBeNull();
// Aucune violation CSP liée à object-src pendant l'ouverture
expect(cspViolations).toEqual([]);
});
});
+19
View File
@@ -302,6 +302,25 @@ async function main() {
}
});
await test("isScrollbarPress: content clicks never unpin, scrollbar drags do (BUG-059)", () => {
const { isScrollbarPress } = bookslmMod;
const container = document.createElement("div");
container.getBoundingClientRect = () => ({
left: 0, right: 800, top: 0, bottom: 600, width: 800, height: 600,
});
const link = document.createElement("a");
// Click on a file link / steps toggle → must NOT unpin (the thread
// would otherwise jump to the bottom when the padding is cleared).
assert.equal(isScrollbarPress(link, 790, container), false);
// Click on blank content area → must NOT unpin either.
assert.equal(isScrollbarPress(container, 300, container), false);
// Dragging the vertical scrollbar: target is the container itself and
// the press lands in the right-edge gutter → unpin allowed.
assert.equal(isScrollbarPress(container, 792, container), true);
// Defensive: no container / no geometry → never unpin.
assert.equal(isScrollbarPress(link, 790, null), false);
});
await test("the placeholder render keeps the anchor (no scroll reset before first token)", async () => {
// Regression: the assistant placeholder used to be rendered while
// `_isLoading` was still false, so it took the "restore previous
+95
View File
@@ -0,0 +1,95 @@
#!/usr/bin/env node
/**
* ObsiGate — Viewer PDF non-regression tests (BUG-060).
*
* Static checks on the source of the PDF inline viewer:
* - BUG-060 : the CSP set by BUG-034 puts `object-src 'none'`, which blocks
* `<embed>`/`<object>`. The PDF body was therefore never rendered (blank
* pages). The viewer must now use an `<iframe>` — permitted by
* `frame-src 'self'` since the PDF stream URL is same-origin.
*
* Usage: node tests/frontend/pdf-viewer.test.mjs
*/
import { strict as assert } from "node:assert";
import { readFileSync } from "node:fs";
import path from "node:path";
import { fileURLToPath } from "node:url";
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const ROOT = path.join(__dirname, "..", "..");
const viewer = readFileSync(path.join(ROOT, "frontend", "js", "viewer.js"), "utf8");
const main = readFileSync(path.join(ROOT, "backend", "main.py"), "utf8");
const css = readFileSync(path.join(ROOT, "frontend", "style.css"), "utf8");
function test(label, fn) {
try {
fn();
console.log(" \u2713 " + label);
} catch (err) {
console.error(" \u2717 " + label + "\n " + err.message);
process.exitCode = 1;
}
}
// ── viewer.js : le rendu PDF passe par une iframe ─────────────────────────
test("viewer.js — PDF branch renders the stream in an <iframe class=pdf-iframe>", () => {
const block = viewer.match(/if \(data\.is_pdf\) \{([\s\S]*?)\n \}/);
assert.ok(block, "PDF render block not found");
assert.match(
block[1],
/<iframe src="\$\{pdfUrl\}" class="pdf-iframe"/,
"PDF must use <iframe>, not <embed>/<object> (CSP object-src 'none' otherwise blocks it)",
);
assert.match(
block[1],
/\.pdf-iframe[\s\S]*?src="\$\{pdfUrl\}"/,
"iframe src must come from the /pdf/stream URL",
);
});
test("viewer.js — no <embed>/<object> left in the source", () => {
assert.doesNotMatch(viewer, /<embed\b/i, "<embed> is blocked by CSP object-src 'none'");
assert.doesNotMatch(viewer, /<object\b/i, "<object> is blocked by CSP object-src 'none'");
});
test("viewer.js — TOC still targets the pdf-iframe via contentWindow", () => {
assert.match(
viewer,
/document\.querySelector\('\.pdf-iframe'\)\.contentWindow\.location\.hash='page=\$\{item\.page\}'/,
"TOC links must keep navigating the iframe",
);
});
// ── backend : la CSP autorise le cadre same-origin ─────────────────────────
test("backend CSP — frame-src 'self' allows same-origin iframes", () => {
const csp = main.match(/frame-src ([^";]+);/);
assert.ok(csp, "CSP frame-src directive not found");
assert.ok(
csp[1].includes("'self'"),
`frame-src must allow 'self' (found: ${csp[1]}) — otherwise the PDF iframe is blocked`,
);
});
test("backend CSP — object-src 'none' stays in place (no defusing)", () => {
assert.match(
main,
/object-src 'none';/,
"object-src must stay locked to 'none': the fix is moving to <iframe>, not weakening CSP",
);
});
// ── style.css : l'iframe garde une hauteur utile ───────────────────────────
test("style.css — .pdf-iframe fills the viewer body", () => {
const rule = css.match(/\.pdf-iframe\s*\{([^}]*)\}/);
assert.ok(rule, ".pdf-iframe rule not found");
assert.match(rule[1], /flex:\s*1/, "iframe must stretch to fill the available height");
assert.match(rule[1], /min-height/, "iframe must keep its minimum height");
});
if (process.exitCode) {
console.error("\nPDF viewer tests FAILED");
} else {
console.log("\nAll PDF viewer tests passed.");
}
+230
View File
@@ -0,0 +1,230 @@
"""Unit tests for the connected sources (#92): Gitea & GitHub tools.
All HTTP calls are mocked (httpx.request monkeypatched) — the CI never talks
to a real Gitea/GitHub instance.
"""
import base64
from typing import Any
import pytest
import backend.tools.connected as connected
from backend.tools.context import ToolContext, ToolError, ToolMode
from backend.tools.registry import get_tool
class FakeResponse:
def __init__(self, json_data: Any = None, status_code: int = 200):
self._json = json_data
self.status_code = status_code
def json(self):
return self._json
def raise_for_status(self):
if self.status_code >= 400:
import httpx
raise httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
@pytest.fixture
def gitea_env(monkeypatch):
monkeypatch.setenv("OBSIGATE_GITEA_URL", "https://git.example.net")
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "tok-gitea")
@pytest.fixture
def github_env(monkeypatch):
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "tok-gh")
def _patch_request(monkeypatch, handler):
def fake_request(method, url, headers=None, timeout=None, follow_redirects=False, **kw):
captured = {"method": method, "url": str(url), "headers": headers or {},
"params": kw.get("params")}
return handler(captured)
monkeypatch.setattr(connected.httpx, "request", fake_request)
class TestRegistration:
@pytest.mark.parametrize("name", ["git_list_repos", "git_search_issues", "git_get_file"])
def test_tools_registered_read(self, name):
from backend.tools.context import ToolRisk
spec = get_tool(name)
assert spec is not None
assert spec.risk == ToolRisk.READ
assert "gitea" in spec.description or "github" in spec.description.lower()
class TestConfiguration:
def test_gitea_requires_base_url(self, monkeypatch):
monkeypatch.delenv("OBSIGATE_GITEA_URL", raising=False)
with pytest.raises(ToolError) as ei:
connected._provider_base("gitea")
assert ei.value.code == "provider_not_configured"
def test_unknown_provider_rejected(self):
with pytest.raises(ToolError) as ei:
connected._provider_base("gitlab")
assert ei.value.code == "invalid_arguments"
def test_gitea_token_sent_as_header(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"data": []})
_patch_request(monkeypatch, handler)
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert captured["headers"]["Authorization"] == "token tok-gitea"
class TestListRepos:
def test_gitea_search_endpoint(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"data": [
{"name": "ObsiGate", "full_name": "bruno/ObsiGate",
"html_url": "https://git.example.net/bruno/ObsiGate",
"description": "vault gateway", "updated_at": "2026-09-01",
"private": False},
]})
_patch_request(monkeypatch, handler)
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert "/api/v1/repos/search" in captured["url"]
assert out["repos"][0]["name"] == "ObsiGate"
assert out["count"] == 1
def test_github_single_repo(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={
"name": "ObsiGate", "full_name": "bruno/ObsiGate",
"html_url": "https://github.com/bruno/ObsiGate",
"description": "", "updated_at": "2026-09-02", "private": True,
})
_patch_request(monkeypatch, handler)
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(
provider="github", repo="bruno/ObsiGate"))
assert captured["url"].endswith("/repos/bruno/ObsiGate")
assert out["repos"][0]["full_name"] == "bruno/ObsiGate"
assert out["repos"][0]["private"] is True
class TestSearchIssues:
def test_gitea_scoped_to_repo(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data=[
{"number": 12, "title": "Bug affichage", "html_url": "https://x/12",
"state": "open"},
])
_patch_request(monkeypatch, handler)
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
provider="gitea", query="affichage", repo="bruno/ObsiGate"))
assert "/repos/bruno/ObsiGate/issues" in captured["url"]
assert out["issues"][0]["id"] == 12
assert out["issues"][0]["pull_request"] is False
def test_github_search_syntax(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"items": [
{"number": 5, "title": "Crash on save", "html_url": "https://gh/5",
"state": "open", "pull_request": {"url": "x"}},
]})
_patch_request(monkeypatch, handler)
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
provider="github", query="crash", repo="bruno/ObsiGate", state="open"))
assert "/search/issues" in captured["url"]
assert "repo:bruno/ObsiGate" in captured["params"]["q"]
assert out["issues"][0]["pull_request"] is True
class TestGetFile:
def test_gitea_base64_content_decoded(self, gitea_env, monkeypatch):
payload = base64.b64encode("# Readme\n\nBonjour".encode()).decode()
def handler(_captured):
return FakeResponse(json_data={
"path": "README.md", "size": 15, "encoding": "base64", "content": payload,
})
_patch_request(monkeypatch, handler)
out = connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="gitea", repo="bruno/ObsiGate", path="README.md"))
assert "Bonjour" in out["content"]
assert out["truncated"] is False
def test_github_ref_parameter(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={
"path": "a.md", "size": 1, "encoding": "base64",
"content": base64.b64encode(b"x").decode(),
})
_patch_request(monkeypatch, handler)
connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="github", repo="o/r", path="a.md", ref="v2.9.0"))
assert captured["url"].endswith("?ref=v2.9.0")
def test_missing_repo_or_path_rejected(self, gitea_env):
with pytest.raises(ToolError) as ei:
connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="gitea", repo="", path="a.md"))
assert ei.value.code == "invalid_arguments"
class TestErrors:
def test_404_maps_to_not_found(self, gitea_env, monkeypatch):
def handler(_captured):
return FakeResponse(json_data={}, status_code=404)
_patch_request(monkeypatch, handler)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "not_found"
def test_401_maps_to_permission_denied(self, gitea_env, monkeypatch):
def handler(_captured):
return FakeResponse(json_data={}, status_code=401)
_patch_request(monkeypatch, handler)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "permission_denied"
def test_network_error_maps_to_tool_error(self, gitea_env, monkeypatch):
import httpx
def fake_request(*a, **kw):
raise httpx.ConnectError("down")
monkeypatch.setattr(connected.httpx, "request", fake_request)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "connected_source_unavailable"
+135
View File
@@ -0,0 +1,135 @@
"""Unit tests for the bounded site crawler (#92): crawl_site (WRITE + confirmation)."""
from typing import Any
import pytest
import backend.tools.crawler as crawler
from backend.tools.api import ToolConfirmationRequired, ToolContext, ToolError, call_tool
class FakeResponse:
def __init__(self, content: bytes = b"", status_code: int = 200,
headers: dict | None = None):
self.content = content
self.status_code = status_code
self.headers = headers or {"content-type": "text/html; charset=utf-8"}
self.encoding = "utf-8"
def raise_for_status(self):
pass
def _ctx() -> ToolContext:
return ToolContext(
user={"username": "tester", "role": "admin", "vaults": ["*"]},
audit_enabled=False,
)
@pytest.fixture
def vault(tmp_path, monkeypatch):
vault_dir = tmp_path / "Vault"
vault_dir.mkdir()
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
return vault_dir
PAGE_A = (
b"<html><head><title>Docs</title></head><body>"
b"<p>Bienvenue sur la documentation.</p>"
b'<a href="/page-b">Suite</a><a href="https://other.dev/x">ext</a>'
b"</body></html>"
)
PAGE_B = (
b"<html><head><title>Page B</title></head><body><p>Details ici.</p></body></html>"
)
@pytest.fixture
def local_urls(monkeypatch):
"""Skip the DNS-based SSRF guard: test hosts are fake, HTTP is mocked."""
monkeypatch.setattr(crawler, "_assert_public_http_url", lambda url: url)
@pytest.fixture
def two_pages(monkeypatch, local_urls):
def fake_get(url, **kw):
url = str(url)
if url.endswith("/page-b"):
return FakeResponse(content=PAGE_B)
return FakeResponse(content=PAGE_A)
monkeypatch.setattr(crawler.httpx, "get", fake_get)
class TestConfirmation:
def test_requires_confirmation(self, vault, two_pages):
with pytest.raises(ToolConfirmationRequired):
call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "Crawls/docs.md",
})
class TestCrawl:
def test_saves_same_host_pages(self, vault, two_pages):
out = call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "Crawls/docs.md",
}, confirm=True)
assert out.ok and out.data["pages"] == 2
digest = (vault / "Crawls" / "docs.md").read_text(encoding="utf-8")
assert "# Crawl de docs.example.dev" in digest
assert "Bienvenue sur la documentation." in digest
assert "Details ici." in digest
assert "other.dev" not in digest
def test_max_pages_bound(self, vault, monkeypatch, local_urls):
# A link farm: every page links to a new page — cap at max_pages.
def fake_get(url, **kw):
url = str(url)
n = int(url.rsplit("/", 1)[-1] or 0)
return FakeResponse(
content=f"<html><head><title>P{n}</title></head><body>"
f"<p>page {n}</p><a href=\"/{n + 1}\">next</a></body></html>".encode())
monkeypatch.setattr(crawler.httpx, "get", fake_get)
out = call_tool("crawl_site", _ctx(), {
"url": "https://farm.example.dev/0",
"vault": "Vault", "path": "farm.md", "max_pages": 3,
}, confirm=True)
assert out.ok and out.data["pages"] == 3
def test_no_pages_recovered(self, vault, monkeypatch, local_urls):
def dead_get(*a, **kw):
raise crawler.httpx.ConnectError("down")
monkeypatch.setattr(crawler.httpx, "get", dead_get)
with pytest.raises(ToolError) as ei:
call_tool("crawl_site", _ctx(), {
"url": "https://dead.example.dev/", "vault": "Vault", "path": "x.md",
}, confirm=True)
assert ei.value.code == "crawl_failed"
def test_internal_url_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("crawl_site", _ctx(), {
"url": "http://127.0.0.1:8080/", "vault": "Vault", "path": "x.md",
}, confirm=True)
assert ei.value.code in ("ssrf_blocked", "dns_error")
def test_binary_content_skipped(self, vault, monkeypatch, local_urls):
def fake_get(url, **kw):
url = str(url)
if url.endswith("/x.pdf"):
return FakeResponse(content=b"%PDF-1.4", headers={"content-type": "application/pdf"})
return FakeResponse(content=PAGE_A)
monkeypatch.setattr(crawler.httpx, "get", fake_get)
out = call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "docs.md", "max_pages": 5,
}, confirm=True)
assert out.ok and out.data["pages"] >= 1
+197
View File
@@ -0,0 +1,197 @@
"""Unit tests for the document-production tools (#92): create_xlsx, create_docx,
create_csv, create_pdf — WRITE risk, confirmation gating, vault persistence."""
import csv
import io
from pathlib import Path
import pytest
from backend.tools.api import (
ToolConfirmationRequired,
ToolContext,
ToolError,
call_tool,
get_tool,
)
from backend.tools.context import ToolRisk
@pytest.fixture
def vault(tmp_path, monkeypatch):
"""A minimal configured vault (index entry patched, no full build)."""
vault_dir = tmp_path / "Vault"
vault_dir.mkdir()
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
return vault_dir
def _ctx() -> ToolContext:
return ToolContext(
user={"username": "tester", "role": "admin", "vaults": ["*"]},
audit_enabled=False,
)
class TestRegistry:
@pytest.mark.parametrize("name", ["create_xlsx", "create_docx", "create_csv", "create_pdf"])
def test_write_risk_and_confirmation(self, name):
spec = get_tool(name)
assert spec is not None
assert spec.risk == ToolRisk.WRITE
assert spec.requires_confirmation is True
class TestConfirmationGating:
def test_csv_requires_confirmation(self, vault):
with pytest.raises(ToolConfirmationRequired):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "data.csv",
"rows": [["a", "b"], [1, 2]],
})
def test_pdf_requires_confirmation(self, vault):
with pytest.raises(ToolConfirmationRequired):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "doc.pdf", "title": "T", "content": "# H\npara",
})
class TestCreateCsv:
def test_creates_file_in_vault(self, vault):
out = call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "Exports/data.csv",
"rows": [["nom", "score"], ["alice", 12], ["bob", 9.5]],
}, confirm=True)
assert out.ok and out.data["success"] is True
path = vault / "Exports" / "data.csv"
assert path.exists()
rows = list(csv.reader(io.StringIO(path.read_text(encoding="utf-8"))))
assert rows[0] == ["nom", "score"]
assert rows[2] == ["bob", "9.5"]
def test_semicolon_delimiter(self, vault):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.csv", "delimiter": ";",
"rows": [["a", "b"], [1, 2]],
}, confirm=True)
content = (vault / "d.csv").read_text(encoding="utf-8")
assert "a;b" in content
def test_wrong_extension_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.txt", "rows": [["a"], [1]],
}, confirm=True)
assert ei.value.code == "invalid_arguments"
def test_empty_rows_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.csv", "rows": [],
}, confirm=True)
class TestCreateXlsx:
def test_creates_readable_workbook(self, vault):
call_tool("create_xlsx", _ctx(), {
"vault": "Vault", "path": "Rapports/budget.xlsx",
"rows": [["item", "cout"], ["serveur", 1200], ["licence", 300]],
"sheet_name": "Budget",
}, confirm=True)
from openpyxl import load_workbook
wb = load_workbook(vault / "Rapports" / "budget.xlsx")
ws = wb.active
assert ws.title == "Budget"
assert ws.cell(row=1, column=1).value == "item"
assert ws.cell(row=2, column=2).value == 1200
def test_wrong_extension_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_xlsx", _ctx(), {
"vault": "Vault", "path": "b.docx", "rows": [["a"], [1]],
}, confirm=True)
class TestCreateDocx:
def test_creates_readable_document(self, vault):
call_tool("create_docx", _ctx(), {
"vault": "Vault", "path": "rapport.docx",
"title": "Rapport hebdo", "paragraphs": ["Premier point.", "Second point."],
}, confirm=True)
import docx as docx_lib
doc = docx_lib.Document(str(vault / "rapport.docx"))
texts = [p.text for p in doc.paragraphs]
assert "Rapport hebdo" in texts
assert "Second point." in texts
def test_no_paragraphs_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_docx", _ctx(), {
"vault": "Vault", "path": "r.docx", "paragraphs": [],
}, confirm=True)
class TestCreatePdf:
def test_creates_valid_pdf(self, vault):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "docs/archi.pdf",
"title": "Architecture", "content": "# Titre\n\nUn paragraphe.\n## Sous-titre\nAutre texte.",
}, confirm=True)
raw = (vault / "docs" / "archi.pdf").read_bytes()
assert raw.startswith(b"%PDF")
def test_long_content_truncated(self, vault):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "big.pdf", "title": "T", "content": "x" * 500_000,
}, confirm=True)
assert (vault / "big.pdf").exists()
def test_markdown_tables_go_through_the_export_pipeline(self, vault, monkeypatch):
# The document-page pipeline (mistune tables + WeasyPrint) must be
# used when available: capture the HTML handed to the PDF generator.
import sys
import types
captured = {}
fake = types.ModuleType("backend.pdf_export")
def fake_build(html, title, **kw):
captured["html"] = html
captured["title"] = title
return "<html>" + html + "</html>"
fake.build_pdf_html = fake_build
fake.generate_pdf = lambda html, title=None, **kw: b"%PDF-fake"
monkeypatch.setitem(sys.modules, "backend.pdf_export", fake)
out = call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "table.pdf", "title": "Rapport",
"content": "# T\n\n| a | b |\n|---|---|\n| 1 | 2 |",
}, confirm=True)
assert out.ok
assert "<table>" in captured["html"]
assert captured["title"] == "Rapport"
assert (vault / "table.pdf").read_bytes() == b"%PDF-fake"
def test_reportlab_fallback_when_weasyprint_missing(self, vault, monkeypatch):
# sys.modules[name] = None makes `from backend.pdf_export import …`
# raise ImportError → the simplified renderer must take over.
import sys
monkeypatch.setitem(sys.modules, "backend.pdf_export", None)
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "fb.pdf", "title": "T", "content": "# H\ntexte",
}, confirm=True)
assert (vault / "fb.pdf").read_bytes().startswith(b"%PDF")
class TestVaultSafety:
def test_path_outside_vault_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "../outside.csv", "rows": [["a"], [1]],
}, confirm=True)
assert ei.value.code in ("path_outside_vault", "invalid_arguments", "tool_execution_error")
+163
View File
@@ -0,0 +1,163 @@
"""Unit tests for the tool/connected-source key store (#103).
Covers ``backend.tools.secrets`` (precedence, masking, whitelist) and the
``/api/config/tool-keys`` endpoints (masked GET, POST, DELETE, admin-only).
"""
import json
from pathlib import Path
import pytest
from backend.tools.secrets import (
TOOL_KEY_NAMES,
delete_tool_key,
get_tool_key,
is_secret_name,
mask_value,
set_tool_key,
)
@pytest.fixture
def key_store(tmp_path, monkeypatch):
"""Isolated key store directory."""
monkeypatch.setenv("OBSIGATE_DATA_DIR", str(tmp_path))
yield tmp_path
def _write_store(tmp_path, data):
(tmp_path / "api_keys.json").write_text(json.dumps(data), encoding="utf-8")
class TestGetToolKey:
def test_stored_value_takes_precedence_over_env(self, key_store, monkeypatch):
_write_store(key_store, {"OBSIGATE_GITHUB_TOKEN": "stored-token"})
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "env-token")
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == "stored-token"
def test_env_fallback_when_not_stored(self, key_store, monkeypatch):
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "env-token")
assert get_tool_key("OBSIGATE_GITEA_TOKEN") == "env-token"
def test_missing_everywhere_returns_empty(self, key_store, monkeypatch):
monkeypatch.delenv("OBSIGATE_GITHUB_TOKEN", raising=False)
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == ""
def test_non_whitelisted_name_uses_env_only(self, key_store, monkeypatch):
_write_store(key_store, {"OTHER_KEY": "stored"})
monkeypatch.setenv("OTHER_KEY", "env")
assert get_tool_key("OTHER_KEY") == "env"
def test_whitelist_covers_expected_names(self):
assert set(TOOL_KEY_NAMES) == {
"OBSIGATE_TAVILY_API_KEY",
"OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY",
"OBSIGATE_EXA_API_KEY",
"OBSIGATE_GITEA_URL",
"OBSIGATE_GITEA_TOKEN",
"OBSIGATE_GITHUB_TOKEN",
}
class TestSetDelete:
def test_set_then_get_roundtrip(self, key_store):
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == "tvly-1234"
def test_set_empty_value_deletes_entry(self, key_store):
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
set_tool_key("OBSIGATE_TAVILY_API_KEY", "")
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == ""
assert "OBSIGATE_TAVILY_API_KEY" not in json.loads(
(key_store / "api_keys.json").read_text(encoding="utf-8")
)
def test_delete_removes_and_reports(self, key_store):
set_tool_key("OBSIGATE_GITEA_URL", "https://git.example.net")
assert delete_tool_key("OBSIGATE_GITEA_URL") is True
assert delete_tool_key("OBSIGATE_GITEA_URL") is False
def test_unknown_name_rejected(self, key_store):
with pytest.raises(ValueError):
set_tool_key("NOT_WHITELISTED", "x")
with pytest.raises(ValueError):
delete_tool_key("NOT_WHITELISTED")
def test_store_keeps_other_entries(self, key_store):
_write_store(key_store, {"DEEPSEEK_API_KEY": "sk-existing"})
set_tool_key("OBSIGATE_EXA_API_KEY", "exa-key")
data = json.loads((key_store / "api_keys.json").read_text(encoding="utf-8"))
assert data["DEEPSEEK_API_KEY"] == "sk-existing"
assert data["OBSIGATE_EXA_API_KEY"] == "exa-key"
class TestMasking:
def test_urls_returned_clear(self):
assert mask_value("OBSIGATE_GITEA_URL", "https://git.example.net") == \
"https://git.example.net"
def test_tokens_masked(self):
masked = mask_value("OBSIGATE_GITHUB_TOKEN", "ghp_abcdefgh1234")
assert masked.startswith("ghp_")
assert "abcdefgh1234" not in masked
assert "..." in masked
def test_short_secret_fully_masked(self):
assert mask_value("OBSIGATE_EXA_API_KEY", "abc") == "***"
def test_secret_detection(self):
assert is_secret_name("OBSIGATE_GITEA_TOKEN")
assert is_secret_name("OBSIGATE_TAVILY_API_KEY")
assert not is_secret_name("OBSIGATE_GITEA_URL")
class TestToolKeysAPI:
def _login(self, admin_client):
resp = admin_client.post(
"/api/auth/login", json={"username": "admin", "password": "chab30"}
)
assert resp.status_code == 200, resp.text
token = resp.json()["access_token"]
return {"Authorization": f"Bearer {token}"}
def test_get_masks_tokens_shows_urls(self, admin_client, key_store):
headers = self._login(admin_client)
_write_store(key_store, {
"OBSIGATE_GITEA_URL": "https://git.example.net",
"OBSIGATE_GITHUB_TOKEN": "ghp_abcdefgh1234",
})
resp = admin_client.get("/api/config/tool-keys", headers=headers)
assert resp.status_code == 200
data = resp.json()
assert data["OBSIGATE_GITEA_URL"] == "https://git.example.net"
assert "abcdefgh1234" not in data["OBSIGATE_GITHUB_TOKEN"]
assert data["OBSIGATE_GITHUB_TOKEN"].startswith("ghp_")
def test_post_roundtrip_then_delete(self, admin_client, key_store):
headers = self._login(admin_client)
resp = admin_client.post(
"/api/config/tool-keys",
headers=headers,
json={"OBSIGATE_EXA_API_KEY": "exa-key-1234"},
)
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
assert get_tool_key("OBSIGATE_EXA_API_KEY") == "exa-key-1234"
resp = admin_client.delete("/api/config/tool-keys/OBSIGATE_EXA_API_KEY", headers=headers)
assert resp.status_code == 200
assert resp.json()["status"] == "deleted"
assert get_tool_key("OBSIGATE_EXA_API_KEY") == ""
def test_post_unknown_name_rejected(self, admin_client, key_store):
headers = self._login(admin_client)
resp = admin_client.post(
"/api/config/tool-keys", headers=headers, json={"MY_SECRET": "x"}
)
assert resp.status_code == 400
def test_requires_admin(self, admin_client, key_store):
resp = admin_client.get("/api/config/tool-keys", headers={})
assert resp.status_code in (401, 403)
+82
View File
@@ -0,0 +1,82 @@
"""Unit tests for the SQLite web cache (backend.tools.webcache, #92)."""
import pytest
from backend.tools import webcache
@pytest.fixture
def cache_enabled(tmp_path, monkeypatch):
"""Enable the cache against an isolated file with a short TTL."""
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
webcache._schema_ready = False
yield
webcache._schema_ready = False
class TestCacheKey:
def test_deterministic_and_payload_sensitive(self):
a = webcache.cache_key("search", {"q": "pizza", "page": 1})
b = webcache.cache_key("search", {"page": 1, "q": "pizza"})
c = webcache.cache_key("search", {"q": "pasta", "page": 1})
d = webcache.cache_key("fetch", {"q": "pizza", "page": 1})
assert a == b
assert a != c
assert a != d
class TestCacheRoundTrip:
def test_set_get_roundtrip(self, cache_enabled):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"results": [1, 2, 3], "provider": "tavily"})
assert webcache.cache_get(key) == {"results": [1, 2, 3], "provider": "tavily"}
def test_miss_returns_none(self, cache_enabled):
assert webcache.cache_get("search:unknown") is None
def test_overwrite_updates_value(self, cache_enabled):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
webcache.cache_set(key, {"v": 2})
assert webcache.cache_get(key) == {"v": 2}
def test_ttl_expiry(self, tmp_path, monkeypatch):
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
webcache._schema_ready = False
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
assert webcache.cache_get(key) == {"v": 1}
import time
time.sleep(0.15)
assert webcache.cache_get(key) is None
def test_disabled_when_ttl_zero(self, cache_enabled, monkeypatch):
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0")
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
assert webcache.cache_get(key) is None
def test_purge_expired(self, cache_enabled, monkeypatch):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
import time
time.sleep(0.15)
assert webcache.purge_expired() >= 1
assert webcache.cache_get(key) is None
def test_corrupt_db_degrades_silently(self, tmp_path, monkeypatch):
# A directory as cache file breaks sqlite3.connect: the cache must
# disable itself instead of breaking the tools.
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
webcache._schema_ready = False
try:
assert webcache.cache_get("search:x") is None
webcache.cache_set("search:x", {"v": 1})
finally:
webcache._schema_ready = False
+155
View File
@@ -0,0 +1,155 @@
"""Unit tests for the keyed web-search providers (#92): Tavily, Brave,
SerpAPI, Exa — plus provider ordering and the transient retry."""
import httpx
import pytest
import backend.tools.web as web
from backend.tools.context import ToolContext, ToolError, ToolMode
class FakeResponse:
def __init__(self, json_data=None, status_code=200, content=b""):
self._json = json_data
self.status_code = status_code
self.content = content
self.encoding = "utf-8"
self.headers = {}
def json(self):
return self._json
def raise_for_status(self):
if self.status_code >= 400:
raise web.httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
def _no_fallback(monkeypatch):
"""Limit the chain to the provider under test (no searxng/ddg/bing noise)."""
monkeypatch.setattr(web, "WEB_FALLBACK_ENABLED", False)
monkeypatch.setattr(web, "SEARXNG_URL", "http://searxng.invalid")
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: (_ for _ in ()).throw(
httpx.ConnectError("offline")))
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: (_ for _ in ()).throw(
httpx.ConnectError("offline")))
class TestKeyedProviderParsers:
def test_tavily_maps_results(self, monkeypatch):
captured = {}
def fake_post(url, json=None, **kw):
captured["url"] = url
captured["payload"] = json
return FakeResponse(json_data={"results": [
{"title": "T", "url": "https://a.dev", "content": "c" * 800},
]})
monkeypatch.setattr(web.httpx, "post", fake_post)
results, engines = web._search_tavily("q", web.WebSearchInput(query="q", max_results=3))
assert results[0]["title"] == "T"
assert len(results[0]["snippet"]) <= 600
assert engines == []
assert captured["payload"]["api_key"] == ""
assert captured["payload"]["max_results"] == 3
def test_brave_maps_results(self, monkeypatch):
captured = {}
def fake_get(url, params=None, headers=None, **kw):
captured["url"] = url
captured["headers"] = headers
return FakeResponse(json_data={"web": {"results": [
{"title": "B", "url": "https://b.dev", "description": "d"},
]}})
monkeypatch.setattr(web.httpx, "get", fake_get)
results, _ = web._search_brave("q", web.WebSearchInput(query="q"))
assert results[0]["title"] == "B"
assert "api.search.brave.com" in str(captured["url"])
def test_serpapi_maps_results(self, monkeypatch):
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
"organic_results": [{"title": "S", "link": "https://s.dev", "snippet": "sn"}],
}))
results, _ = web._search_serpapi("q", web.WebSearchInput(query="q"))
assert results[0]["url"] == "https://s.dev"
def test_exa_maps_results(self, monkeypatch):
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: FakeResponse(json_data={
"results": [{"title": "E", "url": "https://e.dev", "text": "t" * 900}],
}))
results, _ = web._search_exa("q", web.WebSearchInput(query="q"))
assert results[0]["title"] == "E"
assert len(results[0]["snippet"]) <= 600
class TestProviderChain:
def test_no_key_falls_back_to_searxng(self, monkeypatch):
for var in ("OBSIGATE_TAVILY_API_KEY", "OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY", "OBSIGATE_EXA_API_KEY"):
monkeypatch.delenv(var, raising=False)
chain = [name for name, _ in web._provider_chain()]
assert chain[0] == "searxng"
assert "tavily" not in chain
def test_keyed_provider_used_first_when_key_set(self, monkeypatch):
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
chain = [name for name, _ in web._provider_chain()]
assert chain[0] == "tavily"
assert "brave" not in chain # no key → skipped
def test_explicit_order_env(self, monkeypatch):
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
monkeypatch.setenv("OBSIGATE_EXA_API_KEY", "k")
monkeypatch.setenv("OBSIGATE_WEB_PROVIDERS", "exa,unknown,tavily")
chain = [name for name, _ in web._provider_chain()]
assert chain[:2] == ["exa", "tavily"]
def test_search_uses_keyed_provider_first(self, monkeypatch):
_no_fallback(monkeypatch)
monkeypatch.setenv("OBSIGATE_BRAVE_API_KEY", "k")
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
"web": {"results": [{"title": "B", "url": "https://b.dev", "description": "d"}]},
}))
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert out["provider"] == "brave"
assert out["count"] == 1
class TestRetry:
def test_transient_error_retried_then_succeeds(self, monkeypatch):
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
calls = {"n": 0}
def flaky_get(*a, **kw):
calls["n"] += 1
if calls["n"] == 1:
raise httpx.ConnectError("blip")
return FakeResponse(json_data={"results": [
{"title": "A", "url": "https://a.dev", "content": "x"}]})
monkeypatch.setattr(web.httpx, "get", flaky_get)
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert out["provider"] == "searxng"
assert calls["n"] == 2
def test_persistent_error_not_retried_forever(self, monkeypatch):
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
calls = {"n": 0}
def dead_get(*a, **kw):
calls["n"] += 1
raise httpx.ConnectError("down")
monkeypatch.setattr(web.httpx, "get", dead_get)
with pytest.raises(ToolError) as ei:
web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert ei.value.code == "web_search_unavailable"
assert calls["n"] == 2 # initial + 1 retry, per provider
+73
View File
@@ -0,0 +1,73 @@
"""Unit tests for the dynamic rendering path (#92): fetch_url(render=True)."""
import pytest
import backend.tools.web as web
from backend.tools import webrender
from backend.tools.context import ToolContext, ToolError, ToolMode
from backend.tools.registry import get_tool
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
class TestRegistration:
def test_render_param_exposed_in_schema(self):
spec = get_tool("fetch_url")
assert spec is not None
assert "render" in spec.input_model.model_fields
class TestRenderUnavailable:
def test_missing_playwright_clear_error(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: False)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert ei.value.code == "playwright_unavailable"
def test_ssrf_guard_applied_before_render(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="http://127.0.0.1:9222/devtools", render=True))
assert ei.value.code in ("ssrf_blocked", "dns_error")
class TestRenderSuccess:
def test_fetch_url_delegates_to_worker(self, monkeypatch):
captured = {}
def fake_render(url):
captured["url"] = url
return {"url": url, "status": 200, "title": "SPA",
"text": "dynamic content", "rendered": True, "truncated": False}
monkeypatch.setattr(webrender, "render_page", fake_render)
out = web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert captured["url"] == "https://example.com/spa"
assert out["rendered"] is True
assert "dynamic content" in out["text"]
def test_worker_failure_maps_to_tool_error(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
def boom(url):
raise RuntimeError("chromium crashed")
# The executor re-raises the worker exception on .result(); render_page
# must wrap it into a ToolError instead of leaking a bare exception.
monkeypatch.setattr(webrender, "_render_in_worker", boom)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert ei.value.code == "render_unavailable"
class TestMarkdownExtraction:
def test_html_to_text_reused(self):
text = webrender._html_to_text("<html><body><p>hello</p><script>x()</script></body></html>")
assert "hello" in text
assert "x()" not in text