Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7f0f64a42e | ||
|
|
f7e068baed | ||
|
|
2e2a33cef3 | ||
|
|
2c460022f8 | ||
|
|
133644a0ba | ||
|
|
ba0ec3d1fa | ||
|
|
22e9240e4f | ||
|
|
6a58a59a11 | ||
|
|
26328fadeb |
@@ -7,6 +7,11 @@ OBSIGATE_AUTH_ENABLED=true
|
||||
OBSIGATE_ADMIN_USER=admin
|
||||
OBSIGATE_ADMIN_PASSWORD=chab30
|
||||
|
||||
# DANGER : si OBSIGATE_AUTH_ENABLED=false, toute requête devient un admin
|
||||
# anonyme. Le serveur REFUSE de démarrer sur une adresse non-loopback
|
||||
# (ex. 0.0.0.0) sauf si l'on force l'opt-in ci-dessous. À réserver au local.
|
||||
# OBSIGATE_ALLOW_INSECURE=false
|
||||
|
||||
# Sécurité des cookies (activer si derrière HTTPS)
|
||||
# OBSIGATE_SECURE_COOKIES=false
|
||||
|
||||
@@ -74,3 +79,19 @@ DEEPSEEK_MODEL=deepseek-chat
|
||||
# Chaîne de repli sans clé (DuckDuckGo puis Bing) si SearXNG ne remonte rien
|
||||
# OBSIGATE_WEB_FALLBACK=1
|
||||
# OBSIGATE_WEB_TIMEOUT=10
|
||||
# Fournisseurs à clé (#92), essayés avant SearXNG — injecter via Infisical en prod
|
||||
# OBSIGATE_TAVILY_API_KEY=
|
||||
# OBSIGATE_BRAVE_API_KEY=
|
||||
# OBSIGATE_SERPAPI_API_KEY=
|
||||
# OBSIGATE_EXA_API_KEY=
|
||||
# Ordre des fournisseurs (sinon : clés présentes puis SearXNG puis replis)
|
||||
# OBSIGATE_WEB_PROVIDERS=brave,searxng
|
||||
# Réessais réseau (backoff maison) + cache SQLite des résultats web
|
||||
# OBSIGATE_WEB_RETRY=1
|
||||
# OBSIGATE_WEB_CACHE_TTL=900 # secondes ; 0 = cache désactivé
|
||||
# Rendu dynamique (pages SPA) — dépendance optionnelle :
|
||||
# pip install playwright && playwright install chromium
|
||||
# ── Assistant IA — sources connectées (Gitea / GitHub) ──
|
||||
# OBSIGATE_GITEA_URL=https://git.example.net
|
||||
# OBSIGATE_GITEA_TOKEN=
|
||||
# OBSIGATE_GITHUB_TOKEN=
|
||||
|
||||
@@ -38,6 +38,7 @@ jobs:
|
||||
- name: Frontend unit tests
|
||||
run: |
|
||||
node tests/frontend/unit.test.mjs
|
||||
node tests/frontend/pdf-viewer.test.mjs
|
||||
node tests/frontend/forge-completion.test.mjs
|
||||
|
||||
- name: Frontend JSDOM tests (PaneManager + Excalidraw + Plugins + AI + SW + Collab + Mobile + Semantic + Desktop + Inline edition)
|
||||
@@ -196,6 +197,7 @@ jobs:
|
||||
-e DIR_1_NAME=TestDir \
|
||||
-e DIR_1_PATH=/vaults/TestDir \
|
||||
-e OBSIGATE_AUTH_ENABLED=false \
|
||||
-e OBSIGATE_ALLOW_INSECURE=true \
|
||||
obsigate:ci
|
||||
# Docker-in-docker : le bind mount $(pwd)/... pointe sur un chemin
|
||||
# du job container, inexistant sur l'hôte → montage vide. Les -v
|
||||
|
||||
+189
-1
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
|
||||
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
> **En cours de développement** : les changements à venir sont listés dans la section
|
||||
> [Unreleased](#unreleased). La dernière version livrée est **2.9.0**.
|
||||
> [Unreleased](#unreleased). La dernière version livrée est **2.11.5**.
|
||||
|
||||
---
|
||||
|
||||
@@ -14,6 +14,194 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
---
|
||||
|
||||
## [2.11.5] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-063 (complément) — Viewer PDF : la table des matières ne déplaçait toujours pas la
|
||||
page** : le premier correctif réassignait `iframe.src` avec le seul fragment `#page=N`, ce qui
|
||||
ne change que le fragment → navigation *same-document* que le lecteur PDF natif ignore (il
|
||||
n'applique `#page=N` qu'au chargement). Diagnostic en Chrome *headful* par comparaison de
|
||||
captures. `navigatePdfToPage()` ajoute désormais un paramètre de query horodaté
|
||||
(`&_pdfpage=<ts>#page=N`) pour forcer un vrai rechargement de l'iframe. Fichiers :
|
||||
`frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs`, `tests/e2e/pdf-viewer.spec.js`.
|
||||
|
||||
---
|
||||
|
||||
## [2.11.4] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-061 — Assistant IA : le bouton « Plein écran » n'agrandissait plus le panneau** : la
|
||||
largeur du panneau est écrite en style inline par la poignée de redimensionnement (et par la
|
||||
largeur persistée en `localStorage`) ; cet inline l'emportait sur la règle
|
||||
`.bookslm-panel.fullscreen { width: 100vw }`, donc le panneau restait à sa largeur courante.
|
||||
La règle plein écran est désormais prioritaire (`!important`). Fichier : `frontend/style.css`.
|
||||
- **BUG-062 — Viewer PDF : largeur incomplète quand la navigation est masquée** : la règle de
|
||||
colonne de lecture centrée (`.sidebar.hidden … { max-width: 1200px }`) s'appliquait aussi aux
|
||||
viewers plein cadre. Les conteneurs PDF et image sont maintenant exemptés
|
||||
(`:has(.pdf-viewer-container)` / `:has(.image-viewer-container)` → `max-width: none`). Fichier :
|
||||
`frontend/style.css`.
|
||||
- **BUG-063 — Viewer PDF : la table des matières ne naviguait pas** : les liens faisaient
|
||||
`contentWindow.location.hash = 'page=N'`, mais le lecteur PDF natif vit dans une fenêtre
|
||||
`about:blank` et l'affectation n'atteignait jamais le document. Nouveau helper
|
||||
`navigatePdfToPage()` qui recharge l'iframe avec le fragment `#page=N` ; les entrées portent un
|
||||
`data-page` et sont câblées par des écouteurs (plus d'`onclick` inline). Fichiers :
|
||||
`frontend/js/viewer.js`, `frontend/style.css`.
|
||||
- **Tests** : `tests/frontend/ai.test.mjs` (+1), `tests/frontend/pdf-viewer.test.mjs` (TOC, plein
|
||||
largeur), `tests/e2e/pdf-viewer.spec.js` (TOC `#page=N`, largeur, fixture
|
||||
`test_vault/sample-pdf-toc.pdf`).
|
||||
|
||||
---
|
||||
|
||||
## [2.11.3] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-035 — `secret_redactor` : faux positifs sur les hashs hex** : la règle qui masquait
|
||||
tout jeton hexadécimal de 40 à 64 caractères mutilait les hashs git/SHA légitimes des notes.
|
||||
Le masquage des chaînes hexadécimales n'a désormais lieu que si un mot-clé de secret
|
||||
(`secret`, `token`, `key`, `password`, `bearer`…) figure dans les 60 caractères précédents ;
|
||||
un contexte de hash (`commit`, `sha256`, `hash`, `checksum`, `git`, `etag`…) exempte
|
||||
explicitement la chaîne. Fichier : `backend/secret_redactor.py`.
|
||||
- **BUG-036 — Collaboration WebSocket : jeton accepté en query string** : le JWT n'est plus lu
|
||||
depuis `?token=` (URLs journalisées par les proxies et l'historique navigateur). Le cookie
|
||||
HttpOnly `access_token`, envoyé automatiquement par le navigateur lors du handshake
|
||||
same-origin, est le seul transport supporté ; les trames brutes dépassant
|
||||
`MAX_MESSAGE_CHARS` (16 Mio) sont rejetées avant analyse. Fichier : `backend/collab.py`.
|
||||
- **BUG-037 — Mode sans authentification** : au démarrage, un avertissement explicite est
|
||||
journalisé quand `OBSIGATE_AUTH_ENABLED=false`. Le serveur **refuse désormais de démarrer**
|
||||
s'il est lié à une adresse non-loopback sans l'opt-in explicite `OBSIGATE_ALLOW_INSECURE=true`,
|
||||
pour empêcher l'exposition publique d'une instance sans authentification (admin anonyme).
|
||||
Fichiers : `backend/auth/middleware.py`, `backend/main.py`.
|
||||
- **BUG-038 — Argon2 : coût mémoire recalibré** : `memory_cost` passe de 64 Mio à 19 Mio
|
||||
(`m=19456 Kio, t=2, p=1`, recommandation OWASP actuelle) pour supprimer le risque
|
||||
d'épuisement mémoire sous connexions simultanées ; les anciens hachages restent valides et
|
||||
sont migrés automatiquement (`needs_rehash`). Fichier : `backend/auth/password.py`.
|
||||
- **BUG-039 — Énumération de comptes au login** : les comptes inconnus, désactivés, verrouillés
|
||||
et limités par le budget par compte répondent tous un `401 Identifiants invalides` avec un
|
||||
temps équivalent (hachage factice), au lieu d'un `429`/`403` distinctif ; seul le rate-limit
|
||||
par IP (non lié à un compte) conserve le `429`. Fichier : `backend/auth/router.py`.
|
||||
- **BUG-040 — Extraction PDF différée au scan** : `_scan_vault` ne lit plus que les métadonnées
|
||||
des PDF ; l'extraction de texte intégrale (100 kio) est déléguée à `enrich_pdf_texts()`,
|
||||
exécutée après la construction de l'index/inverted index (démarrage) et après chaque
|
||||
réindexation. Un vault contenant de nombreux/gros PDF démarre sans être bloqué ; le texte
|
||||
reste recherchable une fois l'enrichissement terminé. Fichiers : `backend/indexer.py`,
|
||||
`backend/main.py`.
|
||||
- **Tests** : `tests/test_api_main.py` (redactor hex), `tests/test_auth.py` (coût Argon2,
|
||||
garde-fou d'instance non authentifiée), `tests/test_auth_api.py` (login uniforme),
|
||||
`tests/test_collab.py` (jeton query rejeté, trame surdimensionnée), `tests/test_pdf.py`
|
||||
(scan différé + enrichissement).
|
||||
|
||||
---
|
||||
|
||||
## [2.11.2] — 2026-09-17
|
||||
|
||||
### Modifié
|
||||
|
||||
- **Tests E2E PDF (BUG-060)** : le helper d'ouverture de fichier de
|
||||
`tests/e2e/pdf-viewer.spec.js` étend désormais le vault dans l'arborescence s'il est replié
|
||||
(cas d'une instance avec authentification activée) au lieu de supposer le vault déjà déplié.
|
||||
|
||||
---
|
||||
|
||||
## [2.11.1] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-060 — Viewer PDF : l'affichage des pages ne fonctionnait pas** : au clic sur un fichier
|
||||
`.pdf`, seule la barre d'outils « PDF — N pages » s'affichait, le corps restant vide. La CSP
|
||||
durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>`, alors
|
||||
que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était
|
||||
bloqué. Le rendu passe désormais par une `<iframe>` (autorisée par `frame-src 'self'`, le stream
|
||||
`/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` est conservé. Fichiers :
|
||||
`frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau),
|
||||
`tests/e2e/pdf-viewer.spec.js` (nouveau, fixture `test_vault/sample-pdf.pdf`).
|
||||
|
||||
---
|
||||
|
||||
## [2.11.0] — 2026-09-17
|
||||
|
||||
### Ajouté
|
||||
|
||||
- **#103 — Clés des sources connectées configurables depuis la page Configurations** : nouvelle
|
||||
section « 🔗 Sources connectées & recherche » (admin) permettant de saisir ses propres jetons
|
||||
sans toucher aux variables d'environnement : clés de recherche web à clé (Tavily, Brave,
|
||||
SerpAPI, Exa), URL Gitea + token, token GitHub. Stockage dans `data/api_keys.json` (même
|
||||
fichier que les clés des fournisseurs IA) via les endpoints
|
||||
`GET/POST/DELETE /api/config/tool-keys` (GET masque les secrets, URLs en clair ; liste
|
||||
blanche stricte ; admin requis). Les outils lisent la valeur stockée **en priorité** puis
|
||||
l'environnement (Infisical) — `backend/tools/secrets.py`. Fichiers :
|
||||
`backend/tools/{secrets,web,connected}.py`, `backend/main.py`, `frontend/index.html`,
|
||||
`frontend/js/config.js`, `frontend/locales/{fr,en}.json`, `tests/test_tool_keys.py` (nouveau),
|
||||
`tests/test_connected_tools.py`.
|
||||
|
||||
---
|
||||
|
||||
## [2.10.1] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-059 — Assistant IA : un clic dans la conversation faisait sauter le texte au bas de la
|
||||
fenêtre** : dans une conversation ouverte (post ancré en haut), tout clic — lien de fichier,
|
||||
étapes, sélection de texte — dépinait l'ancre via le gestionnaire `mousedown` et retirait le
|
||||
padding d'ancre, bornant le scroll à la nouvelle hauteur max (saut au bas). Le dépintage est
|
||||
désormais réservé aux vrais gestes de scroll : molette, tactile et poignée de scroll
|
||||
uniquement (`isScrollbarPress`). Fichiers : `frontend/js/bookslm.js`,
|
||||
`tests/frontend/ai.test.mjs`.
|
||||
- **#92 — `create_pdf` de l'Assistant : tableaux mal formatés** : l'outil `create_pdf`
|
||||
(génération PDF via l'assistant) utilisait un rendu reportlab simplifié sans support des
|
||||
tableaux. Il passe désormais par le même pipeline que le bouton « Télécharger PDF » de la
|
||||
page document (mistune + plugin `table` + WeasyPrint), avec repli automatique sur le rendu
|
||||
simple si WeasyPrint n'est pas disponible (GTK absent). Fichiers :
|
||||
`backend/tools/documents.py`, `tests/test_document_tools.py`.
|
||||
|
||||
---
|
||||
|
||||
## [2.10.0] — 2026-09-17
|
||||
|
||||
### Ajouté
|
||||
|
||||
- **#92 — Assistant IA : écosystème d'outils phase 2 (web étendu, sources connectées, documents)** :
|
||||
- **Recherche web à clé** : fournisseurs optionnels essayés avant SearXNG — Tavily, Brave
|
||||
Search, SerpAPI (Google) et Exa (`OBSIGATE_TAVILY_API_KEY`, `OBSIGATE_BRAVE_API_KEY`,
|
||||
`OBSIGATE_SERPAPI_API_KEY`, `OBSIGATE_EXA_API_KEY`), avec ordre personnalisable via
|
||||
`OBSIGATE_WEB_PROVIDERS`.
|
||||
- **Transverse** : réessais réseau avec backoff maison (`OBSIGATE_WEB_RETRY`) et cache SQLite
|
||||
des résultats web avec TTL (`OBSIGATE_WEB_CACHE_TTL`, 900 s par défaut, `OBSIGATE_WEB_CACHE_PATH`).
|
||||
- **Rendu dynamique** : `fetch_url(render=True)` délègue les pages SPA à un worker Playwright
|
||||
isolé (dépendance optionnelle, dégradation propre si non installée).
|
||||
- **Crawl de site** : `crawl_site` (WRITE, confirmation) capture jusqu'à 20 pages d'un même
|
||||
hôte et enregistre un condensé Markdown dans un vault.
|
||||
- **Sources connectées** : Gitea (`OBSIGATE_GITEA_URL`/`OBSIGATE_GITEA_TOKEN`) et GitHub
|
||||
(`OBSIGATE_GITHUB_TOKEN`) — `git_list_repos`, `git_search_issues`, `git_get_file` (READ,
|
||||
rate-limités, audités). Les drives cloud (Google Drive / OneDrive) restent orientés serveur
|
||||
MCP externe (#79), conformément à la feuille de route.
|
||||
- **Production de documents** : `create_xlsx` (openpyxl), `create_docx` (python-docx),
|
||||
`create_csv`, `create_pdf` (reportlab) — outils WRITE avec confirmation et sauvegarde
|
||||
dans le vault (backup avant écrasement).
|
||||
- Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests unitaires mockés
|
||||
(httpx). Dépendances ajoutées : `openpyxl`, `python-docx`, `reportlab`.
|
||||
Fichiers : `backend/tools/{webcache,webrender,connected,crawler,documents,web,schemas,labels}.py`,
|
||||
`backend/services/mutations.py`, `tests/test_{web_cache,web_search_providers,webrender,connected_tools,document_tools,crawler}.py`.
|
||||
|
||||
---
|
||||
|
||||
## [2.9.1] — 2026-09-17
|
||||
|
||||
### Corrigé
|
||||
|
||||
- **BUG-058 — Éditeur « Editer » : la barre de numérotation de ligne ne suit pas le thème** :
|
||||
CodeMirror peint son gutter (colonne des numéros de ligne) avec des valeurs claires codées
|
||||
en dur (`#f5f5f5`, bordure `#ddd`), si bien qu'en thème sombre la barre restait gris clair
|
||||
alors que le fond de l'éditeur suivait le thème. Le gutter est désormais dérivé des
|
||||
variables CSS du thème ObsiGate (`color-mix(in srgb, var(--text-primary) …)` pour un fond
|
||||
subtil, `--text-secondary` pour les numéros, `--border` pour la séparation), ce qui le fait
|
||||
suivre tous les thèmes et modes (sombre, clair, contraste élevé, sépia). Fichiers :
|
||||
`frontend/style.css`. Tests : `tests/frontend/editor-inline.test.mjs`.
|
||||
|
||||
---
|
||||
|
||||
## [2.9.0] — 2026-09-17
|
||||
|
||||
### Ajouté
|
||||
|
||||
+13
-3
@@ -4,7 +4,7 @@
|
||||
|
||||
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
|
||||
|
||||
[]()
|
||||
[]()
|
||||
[](https://opensource.org/licenses/MIT)
|
||||
[](https://www.docker.com/)
|
||||
[](https://www.python.org/)
|
||||
@@ -282,6 +282,16 @@ Un compte **admin** connecté voit une icône 🛡️ dans le header : liste, cr
|
||||
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Autoriser les webhooks vers des adresses privées/boucle | `false` |
|
||||
| `OBSIGATE_PDF_MAX_SIZE_MB` | Taille max des PDF extraits (text indexation) | `50` |
|
||||
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | Timeout extraction PDF (secondes) | `30` |
|
||||
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Fournisseurs de recherche web à clé (essayés avant SearXNG) | — |
|
||||
| `OBSIGATE_WEB_PROVIDERS` | Ordre des fournisseurs de recherche (ex. `brave,searxng`) | — |
|
||||
| `OBSIGATE_WEB_RETRY` | Réessais réseau des outils web (backoff maison) | `1` |
|
||||
| `OBSIGATE_WEB_CACHE_TTL` | Durée du cache SQLite des résultats web (secondes, `0` = off) | `900` |
|
||||
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Source connectée Gitea (outil `git_list_repos`…) | — |
|
||||
| `OBSIGATE_GITHUB_TOKEN` | Jeton GitHub (outil `git_list_repos`…) | — |
|
||||
|
||||
> Ces clés peuvent aussi être saisies **depuis l'interface** (menu → Configurations →
|
||||
> « Sources connectées & recherche ») : la valeur saisie est stockée dans `data/api_keys.json`
|
||||
> et prime sur la variable d'environnement.
|
||||
|
||||
### Volume pour la persistance
|
||||
|
||||
@@ -916,8 +926,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
|
||||
|
||||
## 📝 Changelog
|
||||
|
||||
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.9.0).
|
||||
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.11.5).
|
||||
|
||||
---
|
||||
|
||||
*Projet : ObsiGate | Version : 2.9.0 | Dernière mise à jour : Juin 2026*
|
||||
*Projet : ObsiGate | Version : 2.11.5 | Dernière mise à jour : Juin 2026*
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
|
||||
|
||||
[]()
|
||||
[]()
|
||||
[](https://opensource.org/licenses/MIT)
|
||||
[](https://www.docker.com/)
|
||||
[](https://www.python.org/)
|
||||
@@ -320,6 +320,16 @@ When an **admin** account is logged in, a 🛡️ icon appears in the header. Cl
|
||||
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Allow webhooks to private/loopback addresses | `false` |
|
||||
| `OBSIGATE_PDF_MAX_SIZE_MB` | Max PDF size for text extraction | `50` |
|
||||
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | PDF extraction timeout (seconds) | `30` |
|
||||
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Keyed web-search providers (tried before SearXNG) | — |
|
||||
| `OBSIGATE_WEB_PROVIDERS` | Search provider order (e.g. `brave,searxng`) | — |
|
||||
| `OBSIGATE_WEB_RETRY` | Web tools network retries (house-made backoff) | `1` |
|
||||
| `OBSIGATE_WEB_CACHE_TTL` | SQLite cache TTL for web results (seconds, `0` = off) | `900` |
|
||||
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Gitea connected source (`git_list_repos`…) | — |
|
||||
| `OBSIGATE_GITHUB_TOKEN` | GitHub token (`git_list_repos`…) | — |
|
||||
|
||||
> These keys can also be entered **from the UI** (menu → Configurations →
|
||||
> "Connected sources & search"): the stored value goes to `data/api_keys.json`
|
||||
> and takes precedence over the environment variable.
|
||||
|
||||
>All these variables are documented in `.env.example`.
|
||||
|
||||
@@ -1085,8 +1095,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
|
||||
|
||||
## 📝 Changelog
|
||||
|
||||
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.9.0).
|
||||
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.11.5).
|
||||
|
||||
---
|
||||
|
||||
*Project: ObsiGate | Version: 2.9.0 | Last updated: May 2026*
|
||||
*Project: ObsiGate | Version: 2.11.5 | Last updated: May 2026*
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
|
||||
import logging
|
||||
import os
|
||||
import sys
|
||||
|
||||
from fastapi import Depends, HTTPException, Request
|
||||
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
|
||||
@@ -17,6 +18,9 @@ logger = logging.getLogger("obsigate.auth.middleware")
|
||||
|
||||
security = HTTPBearer(auto_error=False)
|
||||
|
||||
#: Hosts considered safe to bind without authentication (loopback only).
|
||||
_LOOPBACK_HOSTS = {"127.0.0.1", "::1", "localhost", "0:0:0:0:0:0:0:1"}
|
||||
|
||||
|
||||
def is_auth_enabled() -> bool:
|
||||
"""Check if authentication is enabled via environment variable.
|
||||
@@ -26,6 +30,34 @@ def is_auth_enabled() -> bool:
|
||||
return os.environ.get("OBSIGATE_AUTH_ENABLED", "true").lower() != "false"
|
||||
|
||||
|
||||
def is_insecure_mode_allowed() -> bool:
|
||||
"""True when the operator explicitly accepts running without auth (BUG-037)."""
|
||||
return os.environ.get("OBSIGATE_ALLOW_INSECURE", "false").lower() in ("1", "true", "yes", "on")
|
||||
|
||||
|
||||
def bind_host_from_argv(argv: list[str] | None = None) -> str | None:
|
||||
"""Extract the ``--host`` value from the process arguments (uvicorn), if any.
|
||||
|
||||
Returns ``None`` when no explicit host is passed (uvicorn then defaults to
|
||||
loopback ``127.0.0.1``).
|
||||
"""
|
||||
args = sys.argv if argv is None else argv
|
||||
for i, arg in enumerate(args):
|
||||
if arg == "--host" and i + 1 < len(args):
|
||||
return args[i + 1]
|
||||
if arg.startswith("--host="):
|
||||
return arg.split("=", 1)[1]
|
||||
return None
|
||||
|
||||
|
||||
def is_loopback_host(host: str | None) -> bool:
|
||||
"""True when *host* is a loopback address (or unset → uvicorn default)."""
|
||||
if not host:
|
||||
return True
|
||||
normalized = host.strip().strip("[]").lower()
|
||||
return normalized in _LOOPBACK_HOSTS
|
||||
|
||||
|
||||
def get_current_user(
|
||||
request: Request,
|
||||
credentials: HTTPAuthorizationCredentials | None = Depends(security),
|
||||
|
||||
@@ -1,14 +1,21 @@
|
||||
# backend/auth/password.py
|
||||
# Argon2id password hashing — OWASP 2024 recommended algorithm.
|
||||
# Parameters: time_cost=2, memory_cost=64MB, parallelism=2
|
||||
# Parameters (BUG-038): time_cost=2, memory_cost=19 MiB, parallelism=1
|
||||
# (OWASP current recommendation for Argon2id). The previous 64 MiB setting
|
||||
# allowed memory exhaustion under concurrent login attempts.
|
||||
|
||||
from argon2 import PasswordHasher
|
||||
from argon2.exceptions import VerificationError, VerifyMismatchError
|
||||
|
||||
#: Argon2id cost parameters (OWASP 2024: m=19456 KiB, t=2, p=1).
|
||||
ARGON2_TIME_COST = 2
|
||||
ARGON2_MEMORY_COST_KIB = 19456 # 19 MiB
|
||||
ARGON2_PARALLELISM = 1
|
||||
|
||||
ph = PasswordHasher(
|
||||
time_cost=2,
|
||||
memory_cost=65536, # 64 MB
|
||||
parallelism=2,
|
||||
time_cost=ARGON2_TIME_COST,
|
||||
memory_cost=ARGON2_MEMORY_COST_KIB,
|
||||
parallelism=ARGON2_PARALLELISM,
|
||||
hash_len=32,
|
||||
salt_len=16,
|
||||
)
|
||||
|
||||
+18
-17
@@ -124,31 +124,32 @@ async def auth_status():
|
||||
async def login(body: LoginRequest, response: Response, request: Request):
|
||||
"""Authenticate a user. Returns access token and sets refresh cookie.
|
||||
|
||||
Implements timing-safe responses to prevent user enumeration:
|
||||
a failed login with an unknown user takes the same time as one
|
||||
with a known user (dummy hash is computed).
|
||||
Implements timing-safe responses to prevent user enumeration: a failed
|
||||
login with an unknown user takes the same time as one with a known user
|
||||
(dummy hash is computed). BUG-039: unknown, inactive, locked and
|
||||
per-account rate-limited accounts all answer the same ``401`` so the HTTP
|
||||
status can never reveal whether an account exists.
|
||||
"""
|
||||
client_ip = get_client_ip(request)
|
||||
|
||||
# IP-based rate limiting (10 failures / 15 min per IP). It is not
|
||||
# account-specific, so a 429 here cannot be used to enumerate accounts.
|
||||
if is_rate_limited(client_ip):
|
||||
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
|
||||
|
||||
user = get_user(body.username)
|
||||
|
||||
if not user:
|
||||
# BUG-039: uniform 401 + equivalent timing for every account-state outcome.
|
||||
if not user or not user.get("active"):
|
||||
# Timing-safe: simulate hash computation to prevent user enumeration
|
||||
hash_password("dummy_timing_protection")
|
||||
raise HTTPException(401, "Identifiants invalides")
|
||||
|
||||
if not user.get("active"):
|
||||
raise HTTPException(403, "Compte désactivé")
|
||||
|
||||
# IP-based rate limiting (10 failures / 15 min per IP)
|
||||
client_ip = get_client_ip(request)
|
||||
if is_rate_limited(client_ip):
|
||||
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
|
||||
|
||||
# BUG-031: per-account budget still applies when the attacker rotates IPs.
|
||||
if is_account_rate_limited(body.username):
|
||||
raise HTTPException(429, "Trop de tentatives sur ce compte (15min)")
|
||||
|
||||
if is_locked(body.username):
|
||||
raise HTTPException(429, "Compte temporairement verrouillé (15min)")
|
||||
# Kept indistinguishable from a wrong password (BUG-039).
|
||||
if is_account_rate_limited(body.username) or is_locked(body.username):
|
||||
hash_password("dummy_timing_protection")
|
||||
raise HTTPException(401, "Identifiants invalides")
|
||||
|
||||
if not verify_password(body.password, user["password_hash"]):
|
||||
attempts = record_login_failure(body.username)
|
||||
|
||||
+14
-4
@@ -44,6 +44,9 @@ MAX_UPDATE_BYTES = 8 * 1024 * 1024
|
||||
#: Taille maximale d'un snapshot texte (protection anti-abus).
|
||||
MAX_TEXT_CHARS = 8 * 1024 * 1024
|
||||
|
||||
#: Taille maximale d'un message brut reçu (protection anti-abus, BUG-036).
|
||||
MAX_MESSAGE_CHARS = 16 * 1024 * 1024
|
||||
|
||||
#: Palette de couleurs attribuées aux utilisateurs (curseurs + avatars).
|
||||
PEER_COLORS = [
|
||||
"#e6194b", "#3cb44b", "#4363d8", "#f58231", "#911eb4",
|
||||
@@ -68,9 +71,13 @@ def authenticate_websocket(websocket: WebSocket) -> dict[str, Any] | None:
|
||||
"""Authenticate a WebSocket connection.
|
||||
|
||||
Mirrors :func:`backend.auth.middleware.get_current_user` but works on the
|
||||
WebSocket scope: the JWT is read from the ``access_token`` cookie (sent
|
||||
automatically by same-origin browsers during the handshake) or, as a
|
||||
fallback, from the ``token`` query parameter.
|
||||
WebSocket scope: the JWT is read from the ``access_token`` cookie, which
|
||||
same-origin browsers send automatically during the handshake.
|
||||
|
||||
BUG-036: the token is **never** accepted from the query string anymore —
|
||||
URLs end up in access logs, proxies and browser history. Browsers cannot
|
||||
set custom headers on a WebSocket handshake, so the HttpOnly cookie set at
|
||||
login is the only supported transport.
|
||||
|
||||
Returns the user dict, or ``None`` if authentication fails.
|
||||
"""
|
||||
@@ -88,7 +95,7 @@ def authenticate_websocket(websocket: WebSocket) -> dict[str, Any] | None:
|
||||
"_token_vaults": ["*"],
|
||||
}
|
||||
|
||||
token = websocket.query_params.get("token") or websocket.cookies.get("access_token")
|
||||
token = websocket.cookies.get("access_token")
|
||||
if not token:
|
||||
return None
|
||||
|
||||
@@ -274,6 +281,9 @@ class CollabManager:
|
||||
|
||||
# -- message handling ---------------------------------------------------
|
||||
async def _on_message(self, room: CollabRoom, client: CollabClient, raw: str) -> None:
|
||||
# BUG-036: drop oversized frames before parsing them.
|
||||
if not isinstance(raw, str) or len(raw) > MAX_MESSAGE_CHARS:
|
||||
return
|
||||
try:
|
||||
message = json.loads(raw)
|
||||
except (ValueError, TypeError):
|
||||
|
||||
+74
-6
@@ -481,12 +481,18 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
|
||||
|
||||
# PDF handling — special path (binary, uses pdf_reader)
|
||||
tags: list[str] = []
|
||||
pdf_text_pending = False
|
||||
if ext == ".pdf":
|
||||
from backend.pdf_reader import extract_pdf_metadata, extract_pdf_text
|
||||
raw = extract_pdf_text(fpath, max_chars=100000)
|
||||
from backend.pdf_reader import extract_pdf_metadata
|
||||
# BUG-040: only the (cheap) metadata is read during the
|
||||
# scan. Full-text extraction is deferred to a background
|
||||
# pass (``enrich_pdf_texts``) so a vault with many/large
|
||||
# PDFs no longer blocks startup and index rebuilds.
|
||||
pdf_meta = extract_pdf_metadata(fpath)
|
||||
title = pdf_meta.get("title") or fpath.stem.replace("-", " ").replace("_", " ")
|
||||
content_preview = raw[:200].strip()
|
||||
raw = ""
|
||||
content_preview = ""
|
||||
pdf_text_pending = True
|
||||
elif ext == ".excalidraw" or fpath.name.lower().endswith(".excalidraw.md"):
|
||||
raw = fpath.read_text(encoding="utf-8", errors="replace")
|
||||
raw = extract_excalidraw_indexable(raw)
|
||||
@@ -510,7 +516,7 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
|
||||
title, post.content
|
||||
)
|
||||
|
||||
files.append({
|
||||
file_info = {
|
||||
"path": str(relative).replace("\\", "/"),
|
||||
"title": title,
|
||||
"tags": tags,
|
||||
@@ -519,7 +525,10 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
|
||||
"size": stat.st_size,
|
||||
"modified": modified,
|
||||
"extension": ext,
|
||||
})
|
||||
}
|
||||
if pdf_text_pending:
|
||||
file_info["pdf_text_pending"] = True
|
||||
files.append(file_info)
|
||||
|
||||
for tag in tags:
|
||||
tag_counts[tag] = tag_counts.get(tag, 0) + 1
|
||||
@@ -535,6 +544,60 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
|
||||
return {"files": files, "tags": tag_counts, "path": vault_path, "paths": paths, "config": {}}
|
||||
|
||||
|
||||
async def enrich_pdf_texts(vault_name: str | None = None) -> int:
|
||||
"""Extract text from PDFs whose extraction was deferred during the scan (BUG-040).
|
||||
|
||||
``_scan_vault`` only reads PDF metadata so a vault with many or large PDFs
|
||||
starts serving immediately. This coroutine runs *after* the index (and the
|
||||
inverted index) is ready, extracts the missing text off the event loop and
|
||||
updates the in-memory entry plus the incremental index hooks.
|
||||
|
||||
Args:
|
||||
vault_name: Restrict the pass to a single vault; ``None`` covers every
|
||||
indexed vault.
|
||||
|
||||
Returns:
|
||||
Number of deferred PDFs whose text extraction was attempted.
|
||||
"""
|
||||
from backend.pdf_reader import extract_pdf_text
|
||||
|
||||
pending: list[tuple[str, dict[str, Any], Path]] = []
|
||||
with _index_lock:
|
||||
for name, vault_data in index.items():
|
||||
if vault_name is not None and name != vault_name:
|
||||
continue
|
||||
vault_root = Path(vault_data.get("path", ""))
|
||||
for file_info in vault_data.get("files", []):
|
||||
if file_info.get("pdf_text_pending"):
|
||||
pending.append((name, file_info, vault_root / file_info["path"]))
|
||||
|
||||
if not pending:
|
||||
return 0
|
||||
|
||||
loop = asyncio.get_running_loop()
|
||||
enriched = 0
|
||||
for name, file_info, file_path in pending:
|
||||
try:
|
||||
raw = await loop.run_in_executor(None, extract_pdf_text, file_path, 100000)
|
||||
except Exception as exc: # pragma: no cover - defensive
|
||||
logger.warning("PDF enrichment failed for %s: %s", file_path, exc)
|
||||
raw = ""
|
||||
file_info["content"] = raw[:SEARCH_CONTENT_LIMIT]
|
||||
file_info["content_preview"] = raw[:200].strip()
|
||||
file_info.pop("pdf_text_pending", None)
|
||||
enriched += 1
|
||||
if _on_index_change:
|
||||
try:
|
||||
_on_index_change("add", name, file_info["path"], file_info)
|
||||
except Exception as exc: # pragma: no cover - defensive
|
||||
logger.warning(
|
||||
"Index hook failed after PDF enrichment for %s: %s", file_path, exc
|
||||
)
|
||||
|
||||
logger.info("PDF enrichment: extracted text for %d deferred PDF(s)", enriched)
|
||||
return enriched
|
||||
|
||||
|
||||
async def build_index(progress_callback=None) -> None:
|
||||
"""Build the full in-memory index for all configured vaults.
|
||||
|
||||
@@ -632,6 +695,8 @@ async def reload_index() -> dict[str, Any]:
|
||||
Dict mapping vault names to their file/tag counts.
|
||||
"""
|
||||
await build_index()
|
||||
# BUG-040: complete the deferred PDF extraction for the rebuilt index.
|
||||
await enrich_pdf_texts()
|
||||
stats = {}
|
||||
for name, data in index.items():
|
||||
stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])}
|
||||
@@ -695,7 +760,10 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]:
|
||||
# Rebuild attachment index for this vault only
|
||||
from backend.attachment_indexer import build_attachment_index
|
||||
await build_attachment_index({vault_name: config})
|
||||
|
||||
|
||||
# BUG-040: complete the deferred PDF extraction for this vault.
|
||||
await enrich_pdf_texts(vault_name)
|
||||
|
||||
stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])}
|
||||
logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags")
|
||||
return stats
|
||||
|
||||
+112
-1
@@ -722,12 +722,56 @@ class SecurityHeadersMiddleware(BaseHTTPMiddleware):
|
||||
return response
|
||||
|
||||
|
||||
def _guard_insecure_auth() -> None:
|
||||
"""Warn or refuse to start when authentication is disabled (BUG-037).
|
||||
|
||||
With ``OBSIGATE_AUTH_ENABLED=false`` every request is served as an
|
||||
anonymous admin. That is convenient for local use but dangerous when the
|
||||
process is reachable from a network. Binding to a non-loopback host
|
||||
without the explicit ``OBSIGATE_ALLOW_INSECURE=true`` opt-in is refused.
|
||||
"""
|
||||
from backend.auth.middleware import (
|
||||
bind_host_from_argv,
|
||||
is_auth_enabled,
|
||||
is_insecure_mode_allowed,
|
||||
is_loopback_host,
|
||||
)
|
||||
|
||||
if is_auth_enabled():
|
||||
return
|
||||
|
||||
if is_insecure_mode_allowed():
|
||||
logger.warning(
|
||||
"Authentication is DISABLED and OBSIGATE_ALLOW_INSECURE=true: every request "
|
||||
"is treated as an anonymous administrator. Do not expose this instance."
|
||||
)
|
||||
return
|
||||
|
||||
host = bind_host_from_argv()
|
||||
if not is_loopback_host(host):
|
||||
raise RuntimeError(
|
||||
"Refusing to start: authentication is disabled (OBSIGATE_AUTH_ENABLED=false) "
|
||||
f"while binding to a non-loopback address ('{host}'). This would expose an "
|
||||
"unauthenticated instance with admin access. Enable authentication, or set "
|
||||
"OBSIGATE_ALLOW_INSECURE=true if you really know what you are doing."
|
||||
)
|
||||
|
||||
logger.warning(
|
||||
"Authentication is DISABLED (OBSIGATE_AUTH_ENABLED=false): every request is "
|
||||
"treated as an anonymous administrator. This is only safe on a trusted, "
|
||||
"loopback-only deployment."
|
||||
)
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
"""Application lifespan: build index on startup, cleanup on shutdown."""
|
||||
global _search_executor, _vault_watcher
|
||||
_search_executor = ThreadPoolExecutor(max_workers=2, thread_name_prefix="search")
|
||||
|
||||
|
||||
# BUG-037: refuse to expose an unauthenticated instance on a public bind.
|
||||
_guard_insecure_auth()
|
||||
|
||||
# Bootstrap admin account if needed
|
||||
bootstrap_admin()
|
||||
|
||||
@@ -748,6 +792,11 @@ async def lifespan(app: FastAPI):
|
||||
# Build the semantic (embedding) index in the same background thread pool.
|
||||
await loop.run_in_executor(_search_executor, init_semantic_index)
|
||||
|
||||
# BUG-040: extract the PDF text deferred during the scan now that the
|
||||
# index and inverted index are queryable (keeps startup non-blocking).
|
||||
from backend.indexer import enrich_pdf_texts
|
||||
await enrich_pdf_texts()
|
||||
|
||||
# Scan for plugins in all vaults
|
||||
logger.info("Scanning for plugins...")
|
||||
from backend.indexer import vault_config
|
||||
@@ -3679,6 +3728,68 @@ async def api_delete_ai_key(provider_env: str, current_user=Depends(require_admi
|
||||
return {"status": "deleted", "key": key_name}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Tool & connected-source keys (#103) — same store as the AI provider keys
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
from backend.tools.secrets import (
|
||||
TOOL_KEY_NAMES as _TOOL_KEY_NAMES,
|
||||
)
|
||||
from backend.tools.secrets import (
|
||||
delete_tool_key as _delete_tool_key,
|
||||
)
|
||||
from backend.tools.secrets import (
|
||||
get_tool_key as _get_tool_key,
|
||||
)
|
||||
from backend.tools.secrets import (
|
||||
mask_value as _mask_tool_value,
|
||||
)
|
||||
from backend.tools.secrets import (
|
||||
set_tool_key as _set_tool_key,
|
||||
)
|
||||
|
||||
|
||||
@app.get("/api/config/tool-keys", response_model=AIKeysResponse)
|
||||
async def api_get_tool_keys(current_user=Depends(require_admin)):
|
||||
"""Return tool/connected-source configuration (tokens masked, URLs clear)."""
|
||||
masked = {}
|
||||
for name in _TOOL_KEY_NAMES:
|
||||
masked[name] = _mask_tool_value(name, _get_tool_key(name))
|
||||
return masked
|
||||
|
||||
|
||||
@app.post("/api/config/tool-keys", response_model=StatusResponse)
|
||||
async def api_set_tool_keys(body: dict = Body(...), current_user=Depends(require_admin)):
|
||||
"""Save tool/connected-source keys.
|
||||
|
||||
Only whitelisted names (``backend.tools.secrets.TOOL_KEY_NAMES``) are
|
||||
accepted: Tavily/Brave/SerpAPI/Exa API keys, Gitea URL + token, GitHub
|
||||
token. Empty values delete the stored entry.
|
||||
"""
|
||||
updated = []
|
||||
for name, value in body.items():
|
||||
if name not in _TOOL_KEY_NAMES:
|
||||
raise HTTPException(status_code=400, detail=f"Clé inconnue: {name}")
|
||||
if value is not None and not isinstance(value, str):
|
||||
raise HTTPException(status_code=400, detail=f"Type invalide pour {name}")
|
||||
_set_tool_key(name, value or "")
|
||||
updated.append(name)
|
||||
logger.info(f"Tool keys updated: {updated}")
|
||||
return {"status": "ok"}
|
||||
|
||||
|
||||
@app.delete("/api/config/tool-keys/{name}", response_model=AIKeyDeleteResponse)
|
||||
async def api_delete_tool_key(name: str, current_user=Depends(require_admin)):
|
||||
"""Delete a stored tool key (the environment fallback still applies)."""
|
||||
key_name = name.upper()
|
||||
try:
|
||||
existed = _delete_tool_key(key_name)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
logger.info(f"Tool key deleted: {key_name} (existed={existed})")
|
||||
return {"status": "deleted", "key": key_name}
|
||||
|
||||
|
||||
@app.post("/api/config/ai-keys/test", response_model=AITestResponse)
|
||||
async def api_test_ai_keys(current_user=Depends(require_admin)):
|
||||
"""Test which AI providers are configured.
|
||||
|
||||
@@ -20,3 +20,6 @@ psutil>=5.9
|
||||
pywebpush>=2.3.0
|
||||
mcp==1.9.4
|
||||
sse-starlette==2.1.3
|
||||
openpyxl>=3.1
|
||||
python-docx>=1.1
|
||||
reportlab>=4.0
|
||||
|
||||
@@ -34,7 +34,7 @@ _PATTERNS = [
|
||||
(re.compile(r'(?:api[_-]?key|apikey|secret|token|password|passwd|auth[_-]?token)\s*[:=]\s*[\'"]?([^\s\'"]{20,})[\'"]?', re.IGNORECASE),
|
||||
lambda m: f'{m.group(0).split("=")[0].split(":")[0]}=[MASQUÉ]' if "=" in m.group(0) or ":" in m.group(0) else '[MASQUÉ]'),
|
||||
|
||||
# Generic long hex/base64 strings that look like secrets (40+ chars)
|
||||
# Prefixed API keys (sk-..., pk-..., rk-...)
|
||||
(re.compile(r'(?:sk|pk|rk)-[a-zA-Z0-9]{20,}'), '[CLÉ API MASQUÉE]'),
|
||||
|
||||
# AWS access keys
|
||||
@@ -43,10 +43,50 @@ _PATTERNS = [
|
||||
# GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_)
|
||||
(re.compile(r'gh[pousr]_[a-zA-Z0-9]{36,}'), '[GITHUB_TOKEN MASQUÉ]'),
|
||||
|
||||
# Generic long random-looking strings (40+ hex chars)
|
||||
(re.compile(r'\b[a-fA-F0-9]{40,64}\b'), '[HEX_KEY MASQUÉ]'),
|
||||
]
|
||||
|
||||
# BUG-035: bare 40–64 char hex strings used to be redacted unconditionally,
|
||||
# which mangled legitimate git commit SHAs, checksums and hashes in notes.
|
||||
# They are now only redacted when a secret-ish keyword sits in the immediate
|
||||
# context; hash/commit keywords explicitly exempt them.
|
||||
_HEX_RE = re.compile(r'\b[a-fA-F0-9]{40,64}\b')
|
||||
_SECRET_CONTEXT_RE = re.compile(
|
||||
r'(?i)\b(?:secret|token|key|apikey|api[_-]?key|password|passwd|auth|bearer|'
|
||||
r'credential|x-api-key|x-auth-token)\b'
|
||||
)
|
||||
_HASH_CONTEXT_RE = re.compile(
|
||||
r'(?i)\b(?:commit|sha\d*|hash|md5|blob|git|checksum|digest|integrity|'
|
||||
r'revision|rev|etag|fingerprint)\b'
|
||||
)
|
||||
#: How far before the hex string a keyword may appear to count as context.
|
||||
_HEX_CONTEXT_WINDOW = 60
|
||||
|
||||
|
||||
def _redact_bare_hex_secrets(text: str) -> tuple:
|
||||
"""Redact 40–64 char hex strings only when a secret keyword is nearby.
|
||||
|
||||
Git/SHA/checksum contexts are left untouched (BUG-035).
|
||||
|
||||
Args:
|
||||
text: Text to scan.
|
||||
|
||||
Returns:
|
||||
(redacted_text, redaction_count) tuple.
|
||||
"""
|
||||
count = 0
|
||||
|
||||
def _replace(match: re.Match) -> str:
|
||||
nonlocal count
|
||||
window = text[max(0, match.start() - _HEX_CONTEXT_WINDOW):match.start()]
|
||||
if _HASH_CONTEXT_RE.search(window):
|
||||
return match.group(0)
|
||||
if _SECRET_CONTEXT_RE.search(window):
|
||||
count += 1
|
||||
return '[HEX_KEY MASQUÉ]'
|
||||
return match.group(0)
|
||||
|
||||
return _HEX_RE.sub(_replace, text), count
|
||||
|
||||
|
||||
def redact(text: str) -> tuple:
|
||||
"""Redact sensitive patterns from text.
|
||||
@@ -66,6 +106,8 @@ def redact(text: str) -> tuple:
|
||||
new_result, n = pattern.subn(str(replacement), result)
|
||||
count += n
|
||||
result = new_result
|
||||
result, hex_count = _redact_bare_hex_secrets(result)
|
||||
count += hex_count
|
||||
if count > 0:
|
||||
logger.info(f"Redacted {count} secret(s) from content")
|
||||
return result, count
|
||||
|
||||
@@ -44,7 +44,7 @@ def _ensure_writable(root: Path) -> None:
|
||||
raise ServiceError("Vault is read-only", code="read_only", status=403)
|
||||
|
||||
|
||||
def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
|
||||
def _validate_extension(file_path: Path, *, allow_images: bool = False, allow_docs: bool = False) -> None:
|
||||
"""Reject unsupported file extensions (400)."""
|
||||
from backend.indexer import SUPPORTED_EXTENSIONS
|
||||
|
||||
@@ -53,6 +53,9 @@ def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
|
||||
if allow_images:
|
||||
from backend.attachment_indexer import IMAGE_EXTENSIONS
|
||||
allowed = allowed | IMAGE_EXTENSIONS
|
||||
if allow_docs:
|
||||
# Office documents produced by the AI tool layer (#92).
|
||||
allowed = allowed | {".xlsx", ".docx"}
|
||||
|
||||
if ext not in allowed and file_path.name.lower() not in ("dockerfile", "makefile"):
|
||||
raise ServiceError(
|
||||
@@ -715,17 +718,20 @@ def save_raw_file(
|
||||
content: bytes,
|
||||
*,
|
||||
overwrite: bool = True,
|
||||
allow_docs: bool = False,
|
||||
) -> dict[str, Any]:
|
||||
"""Save a binary or text file to a vault (e.g. from upload / drag-and-drop).
|
||||
|
||||
Creates parent directories automatically and safely validates the path.
|
||||
Supports supported text extensions, images and Excalidraw files.
|
||||
Supports supported text extensions, images, Excalidraw files and — with
|
||||
``allow_docs`` — Office documents (.xlsx/.docx) produced by the AI tools.
|
||||
|
||||
Args:
|
||||
vault_name: Name of the vault.
|
||||
path: Vault-relative path.
|
||||
content: Raw bytes to write.
|
||||
overwrite: When True, replace existing files (with backup).
|
||||
allow_docs: Also accept .xlsx/.docx extensions (AI document tools).
|
||||
|
||||
Returns:
|
||||
Dict with ``success``, ``vault``, ``path``, and ``size``.
|
||||
@@ -733,7 +739,7 @@ def save_raw_file(
|
||||
root = get_vault_root(vault_name)
|
||||
_ensure_writable(root)
|
||||
file_path = resolve_safe_path(root, path)
|
||||
_validate_extension(file_path, allow_images=True)
|
||||
_validate_extension(file_path, allow_images=True, allow_docs=allow_docs)
|
||||
|
||||
rel_path = _rel(root, file_path)
|
||||
|
||||
|
||||
@@ -9,6 +9,9 @@ Note: ObsiGate uses implicit namespace packages (no tracked ``__init__.py``,
|
||||
which ``.gitignore`` excludes via ``_*.py``), hence this explicit facade.
|
||||
"""
|
||||
|
||||
from backend.tools import connected as _connected # noqa: F401 (registers connected-source tools)
|
||||
from backend.tools import crawler as _crawler # noqa: F401 (registers the site crawler)
|
||||
from backend.tools import documents as _documents # noqa: F401 (registers document tools)
|
||||
from backend.tools import service as _service # noqa: F401 (registers tools)
|
||||
from backend.tools import web as _web # noqa: F401 (registers web tools)
|
||||
from backend.tools.context import (
|
||||
|
||||
@@ -0,0 +1,241 @@
|
||||
"""Connected sources — Gitea & GitHub repositories (phase 2 #92).
|
||||
|
||||
The assistant can query the source-hosting platforms the project actually
|
||||
uses (ObsiGate is hosted on Gitea): repositories, issues/pull requests and
|
||||
repository files. Everything is READ-risk, rate-limited through the shared
|
||||
registry and audited.
|
||||
|
||||
Configuration (environment — injected by Infisical in production, never
|
||||
hard-coded):
|
||||
|
||||
* ``OBSIGATE_GITEA_URL`` — base URL of the self-hosted instance (e.g.
|
||||
``https://git.example.net``); the ``gitea`` provider is only available when
|
||||
this variable is set. Admin-controlled, so the SSRF guard does not apply
|
||||
(unlike user-supplied URLs). Both the URL and the tokens can also be set
|
||||
from the configuration page (stored in ``data/api_keys.json``, #103) —
|
||||
the stored value takes precedence over the environment.
|
||||
* ``OBSIGATE_GITEA_TOKEN`` — optional personal access token (private repos).
|
||||
* ``OBSIGATE_GITHUB_TOKEN`` — optional token (raises the API rate limits and
|
||||
unlocks private repositories).
|
||||
|
||||
Cloud drives (Google Drive / OneDrive) deliberately stay out of the core:
|
||||
per the documented roadmap they are best served by an *external MCP server*
|
||||
(#79) so the OAuth surface remains outside ObsiGate.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import binascii
|
||||
import logging
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
|
||||
from backend.tools.context import ToolError, ToolRisk
|
||||
from backend.tools.registry import tool
|
||||
from backend.tools.schemas import GitGetFileInput, GitProviderInput, GitSearchIssuesInput
|
||||
from backend.tools.secrets import get_tool_key
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.connected")
|
||||
|
||||
TIMEOUT = 10.0
|
||||
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
|
||||
MAX_FILE_BYTES = 300_000
|
||||
GITHUB_API = "https://api.github.com"
|
||||
|
||||
|
||||
def _provider_base(provider: str) -> tuple[str, str]:
|
||||
"""Return (base_url, auth_header_value) for the requested provider."""
|
||||
if provider == "gitea":
|
||||
base = get_tool_key("OBSIGATE_GITEA_URL").rstrip("/")
|
||||
if not base:
|
||||
raise ToolError(
|
||||
"Source Gitea non configurée (OBSIGATE_GITEA_URL absente).",
|
||||
code="provider_not_configured",
|
||||
)
|
||||
token = get_tool_key("OBSIGATE_GITEA_TOKEN")
|
||||
return base, f"token {token}" if token else ""
|
||||
if provider == "github":
|
||||
token = get_tool_key("OBSIGATE_GITHUB_TOKEN")
|
||||
return GITHUB_API, f"Bearer {token}" if token else ""
|
||||
raise ToolError(
|
||||
f"Fournisseur inconnu : {provider} ('gitea' ou 'github')",
|
||||
code="invalid_arguments",
|
||||
)
|
||||
|
||||
|
||||
def _headers(auth: str) -> dict[str, str]:
|
||||
headers = {"User-Agent": USER_AGENT, "Accept": "application/json"}
|
||||
if auth:
|
||||
headers["Authorization"] = auth
|
||||
return headers
|
||||
|
||||
|
||||
def _request(method: str, url: str, auth: str, **kwargs: Any) -> httpx.Response:
|
||||
try:
|
||||
resp = httpx.request(
|
||||
method, url, headers=_headers(auth), timeout=TIMEOUT, follow_redirects=False,
|
||||
**kwargs,
|
||||
)
|
||||
except httpx.HTTPError as e:
|
||||
logger.warning("connected source request failed %s: %s", url, e)
|
||||
raise ToolError(
|
||||
"Source connectée momentanément indisponible.",
|
||||
code="connected_source_unavailable",
|
||||
) from e
|
||||
if resp.status_code in (401, 403):
|
||||
raise ToolError(
|
||||
"Accès refusé par la source connectée (jeton manquant ou expiré).",
|
||||
code="permission_denied",
|
||||
)
|
||||
if resp.status_code == 404:
|
||||
raise ToolError("Ressource introuvable sur la source connectée.", code="not_found")
|
||||
resp.raise_for_status()
|
||||
return resp
|
||||
|
||||
|
||||
def _normalize_repo(item: dict[str, Any]) -> dict[str, Any]:
|
||||
return {
|
||||
"name": item.get("name") or "",
|
||||
"full_name": item.get("full_name") or "",
|
||||
"url": item.get("html_url") or item.get("clone_url") or "",
|
||||
"description": item.get("description") or "",
|
||||
"updated": item.get("updated_at") or "",
|
||||
"private": bool(item.get("private", False)),
|
||||
}
|
||||
|
||||
|
||||
@tool(
|
||||
name="git_list_repos",
|
||||
description=(
|
||||
"List repositories on the connected Gitea instance or GitHub account "
|
||||
"(name, url, description, last update). Use when the user asks about "
|
||||
"their code projects."
|
||||
),
|
||||
input_model=GitProviderInput,
|
||||
risk=ToolRisk.READ,
|
||||
)
|
||||
def git_list_repos(ctx, params: GitProviderInput) -> dict[str, Any]:
|
||||
"""Query the configured source and return normalized repositories."""
|
||||
base, auth = _provider_base(params.provider)
|
||||
if params.provider == "gitea":
|
||||
url = base + "/api/v1/repos/search"
|
||||
query: dict[str, Any] = {"limit": params.limit}
|
||||
if params.repo:
|
||||
query["q"] = params.repo
|
||||
resp = _request("GET", url, auth, params=query)
|
||||
items = resp.json().get("data") or []
|
||||
else:
|
||||
if params.repo:
|
||||
url = GITHUB_API + f"/repos/{params.repo.strip('/')}"
|
||||
items = [_request("GET", url, auth).json()]
|
||||
else:
|
||||
resp = _request(
|
||||
"GET", GITHUB_API + "/user/repos",
|
||||
auth, params={"per_page": params.limit, "sort": "updated"},
|
||||
)
|
||||
items = resp.json()
|
||||
repos = [_normalize_repo(item) for item in items if isinstance(item, dict)]
|
||||
return {"provider": params.provider, "count": len(repos), "repos": repos}
|
||||
|
||||
|
||||
@tool(
|
||||
name="git_search_issues",
|
||||
description=(
|
||||
"Search issues and pull requests on the connected Gitea instance or "
|
||||
"GitHub (title/body keywords, optional repository scope, open/closed)."
|
||||
),
|
||||
input_model=GitSearchIssuesInput,
|
||||
risk=ToolRisk.READ,
|
||||
)
|
||||
def git_search_issues(ctx, params: GitSearchIssuesInput) -> dict[str, Any]:
|
||||
"""Query issues (and PRs) from the configured source."""
|
||||
base, auth = _provider_base(params.provider)
|
||||
state = params.state if params.state in ("open", "closed") else "open"
|
||||
if params.provider == "gitea":
|
||||
if params.repo:
|
||||
url = base + f"/api/v1/repos/{params.repo.strip('/')}/issues"
|
||||
query: dict[str, Any] = {"state": state, "limit": params.limit, "q": params.query}
|
||||
resp = _request("GET", url, auth, params=query)
|
||||
items = resp.json()
|
||||
else:
|
||||
url = base + "/api/v1/repos/issues/search"
|
||||
resp = _request("GET", url, auth, params={
|
||||
"q": params.query, "state": state, "limit": params.limit,
|
||||
})
|
||||
items = resp.json()
|
||||
else:
|
||||
clause = f"{params.query} is:issue is:{state}"
|
||||
if params.repo:
|
||||
clause += f" repo:{params.repo.strip('/')}"
|
||||
resp = _request(
|
||||
"GET", GITHUB_API + "/search/issues", auth,
|
||||
params={"q": clause, "per_page": params.limit},
|
||||
)
|
||||
items = (resp.json().get("items") or [])
|
||||
issues = [
|
||||
{
|
||||
"id": item.get("number") or item.get("id") or "",
|
||||
"title": (item.get("title") or "")[:300],
|
||||
"url": item.get("html_url") or "",
|
||||
"state": item.get("state") or "",
|
||||
"pull_request": bool(item.get("pull_request")),
|
||||
}
|
||||
for item in (items if isinstance(items, list) else [])
|
||||
if isinstance(item, dict)
|
||||
]
|
||||
return {
|
||||
"provider": params.provider,
|
||||
"query": params.query,
|
||||
"count": len(issues),
|
||||
"issues": issues,
|
||||
}
|
||||
|
||||
|
||||
@tool(
|
||||
name="git_get_file",
|
||||
description=(
|
||||
"Read a file's content from a connected Gitea or GitHub repository "
|
||||
"(source code, docs, config). Text/JSON only, size-capped."
|
||||
),
|
||||
input_model=GitGetFileInput,
|
||||
risk=ToolRisk.READ,
|
||||
)
|
||||
def git_get_file(ctx, params: GitGetFileInput) -> dict[str, Any]:
|
||||
"""Fetch one repository file and return its decoded text content."""
|
||||
base, auth = _provider_base(params.provider)
|
||||
repo = params.repo.strip("/")
|
||||
path = params.path.strip("/")
|
||||
if not repo or not path:
|
||||
raise ToolError(
|
||||
"'repo' (owner/nom) et 'path' sont obligatoires", code="invalid_arguments"
|
||||
)
|
||||
if params.provider == "gitea":
|
||||
url = base + f"/api/v1/repos/{repo}/contents/{path}"
|
||||
else:
|
||||
url = GITHUB_API + f"/repos/{repo}/contents/{path}"
|
||||
if params.ref:
|
||||
url += f"?ref={params.ref}"
|
||||
resp = _request("GET", url, auth)
|
||||
data = resp.json()
|
||||
encoded = data.get("content") or ""
|
||||
if (data.get("encoding") or "") == "base64" and encoded:
|
||||
try:
|
||||
content = base64.b64decode(encoded).decode("utf-8", errors="replace")
|
||||
except (ValueError, binascii.Error) as e:
|
||||
raise ToolError(
|
||||
"Contenu du fichier illisible (encodage inattendu).",
|
||||
code="file_decode_error",
|
||||
) from e
|
||||
else:
|
||||
content = encoded
|
||||
truncated = len(content) > MAX_FILE_BYTES
|
||||
return {
|
||||
"provider": params.provider,
|
||||
"repo": repo,
|
||||
"path": data.get("path") or path,
|
||||
"size": data.get("size") or len(content),
|
||||
"content": content[:MAX_FILE_BYTES],
|
||||
"truncated": truncated,
|
||||
}
|
||||
@@ -0,0 +1,197 @@
|
||||
"""Multi-page site crawl — ``crawl_site`` (phase 2 #92, WRITE + confirmation).
|
||||
|
||||
The assistant can digest a small public site (documentation, docs portal) and
|
||||
store a Markdown summary inside a vault: one section per page, title, URL and
|
||||
readable text. The crawl is bounded and same-host only:
|
||||
|
||||
* max 20 pages (``max_pages``), same hostname, breadth-first from the entry URL;
|
||||
* SSRF guard on every URL (scheme + private-address rejection), size caps;
|
||||
* no third-party crawler dependency (scrapy deliberately avoided — a bounded
|
||||
httpx BFS keeps the surface small and the runtime predictable; the task is
|
||||
executed as a single background-style tool run instead of a web request
|
||||
pipeline).
|
||||
|
||||
Risk is WRITE: the digest is written into a vault, so the two-step
|
||||
confirmation applies (Apply card in the UI, propose/apply over MCP).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import re
|
||||
import time
|
||||
from typing import Any
|
||||
from urllib.parse import urljoin, urlparse
|
||||
|
||||
import httpx
|
||||
|
||||
from backend.tools.context import ToolContext, ToolError, ToolRisk
|
||||
from backend.tools.registry import tool
|
||||
from backend.tools.schemas import CrawlSiteInput
|
||||
from backend.tools.web import (
|
||||
USER_AGENT,
|
||||
_assert_public_http_url,
|
||||
_html_to_text,
|
||||
_response_text,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.crawler")
|
||||
|
||||
MAX_PAGE_BYTES = 800_000
|
||||
MAX_TOTAL_BYTES = 6_000_000
|
||||
MAX_TEXT_PER_PAGE = 12_000
|
||||
PAGE_TIMEOUT = 10.0
|
||||
_LINK_RE = re.compile(r'<a[^>]*href="([^"#]+)"', re.IGNORECASE)
|
||||
_TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
|
||||
|
||||
|
||||
def _same_host(url: str, host: str) -> bool:
|
||||
return (urlparse(url).hostname or "") == host
|
||||
|
||||
|
||||
def _extract_links(raw: str, base_url: str) -> list[str]:
|
||||
import html as html_lib
|
||||
|
||||
links: list[str] = []
|
||||
for match in _LINK_RE.finditer(raw):
|
||||
href = html_lib.unescape(match.group(1)).strip()
|
||||
if not href or href.lower().startswith(("javascript:", "mailto:", "tel:")):
|
||||
continue
|
||||
absolute = urljoin(base_url, href)
|
||||
if absolute.lower().endswith((".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp", ".pdf", ".zip")):
|
||||
continue
|
||||
links.append(absolute.split("#", 1)[0])
|
||||
return links
|
||||
|
||||
|
||||
def _fetch_page(url: str) -> tuple[str, str]:
|
||||
"""Fetch one page (SSRF-guarded, manual redirects) → (title, text)."""
|
||||
current = _assert_public_http_url(url)
|
||||
resp = None
|
||||
for _hop in range(5):
|
||||
resp = httpx.get(
|
||||
current,
|
||||
headers={"User-Agent": USER_AGENT, "Accept": "text/html,*/*"},
|
||||
timeout=PAGE_TIMEOUT,
|
||||
follow_redirects=False,
|
||||
)
|
||||
if resp.status_code in (301, 302, 303, 307, 308):
|
||||
location = resp.headers.get("location") or ""
|
||||
if not location:
|
||||
break
|
||||
current = _assert_public_http_url(str(httpx.URL(current).join(location)))
|
||||
continue
|
||||
break
|
||||
assert resp is not None
|
||||
resp.raise_for_status()
|
||||
ctype = (resp.headers.get("content-type") or "").lower()
|
||||
if "html" not in ctype and "text" not in ctype:
|
||||
raise ToolError(
|
||||
f"Type de contenu non pris en charge: {ctype.split(';')[0] or 'inconnu'}",
|
||||
code="unsupported_content_type",
|
||||
)
|
||||
raw = (resp.content[:MAX_PAGE_BYTES]).decode(resp.encoding or "utf-8", errors="replace")
|
||||
title_match = _TITLE_RE.search(raw)
|
||||
import html as html_lib
|
||||
|
||||
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
|
||||
return title, _html_to_text(raw)[:MAX_TEXT_PER_PAGE]
|
||||
|
||||
|
||||
@tool(
|
||||
name="crawl_site",
|
||||
description=(
|
||||
"Crawl a small public site (same-host only, max 20 pages) starting at "
|
||||
"a URL and save a Markdown digest (title, url, readable text per page) "
|
||||
"into a vault. Use to capture an online documentation for offline use."
|
||||
),
|
||||
input_model=CrawlSiteInput,
|
||||
risk=ToolRisk.WRITE,
|
||||
requires_vault=True,
|
||||
)
|
||||
def crawl_site(ctx: ToolContext, params: CrawlSiteInput) -> dict[str, Any]:
|
||||
"""Bounded BFS crawl; writes the digest file and returns a summary."""
|
||||
from backend.services.errors import ServiceError
|
||||
from backend.services.mutations import save_raw_file
|
||||
|
||||
start = _assert_public_http_url(params.url.strip())
|
||||
host = urlparse(start).hostname or ""
|
||||
if not host:
|
||||
raise ToolError("URL sans hôte", code="invalid_url")
|
||||
|
||||
queue: list[str] = [start]
|
||||
seen: set[str] = {start}
|
||||
pages: list[dict[str, Any]] = []
|
||||
total_bytes = 0
|
||||
failures: list[str] = []
|
||||
|
||||
while queue and len(pages) < params.max_pages and total_bytes < MAX_TOTAL_BYTES:
|
||||
url = queue.pop(0)
|
||||
try:
|
||||
title, text = _fetch_page(url)
|
||||
except ToolError as e:
|
||||
failures.append(url)
|
||||
logger.warning("crawl_site page failed %s: %s", url, e.code)
|
||||
continue
|
||||
except httpx.HTTPError as e:
|
||||
failures.append(url)
|
||||
logger.warning("crawl_site page failed %s: %s", url, e)
|
||||
continue
|
||||
pages.append({"url": url, "title": title, "text": text})
|
||||
total_bytes += len(text)
|
||||
if len(pages) >= params.max_pages:
|
||||
break
|
||||
try:
|
||||
raw_resp = httpx.get(
|
||||
url, headers={"User-Agent": USER_AGENT}, timeout=PAGE_TIMEOUT,
|
||||
follow_redirects=False,
|
||||
)
|
||||
raw = _response_text(raw_resp)
|
||||
except (httpx.HTTPError, ValueError):
|
||||
continue
|
||||
for link in _extract_links(raw, url):
|
||||
if len(pages) + len(queue) >= params.max_pages:
|
||||
break
|
||||
if link in seen or not _same_host(link, host):
|
||||
continue
|
||||
try:
|
||||
_assert_public_http_url(link)
|
||||
except ToolError:
|
||||
continue
|
||||
seen.add(link)
|
||||
queue.append(link)
|
||||
|
||||
if not pages:
|
||||
raise ToolError(
|
||||
"Aucune page n'a pu être récupérée pour ce site.",
|
||||
code="crawl_failed",
|
||||
)
|
||||
|
||||
lines = [
|
||||
f"# Crawl de {host}",
|
||||
"",
|
||||
f"> {len(pages)} page(s) capturée(s) depuis {start} — {time.strftime('%Y-%m-%d %H:%M')}",
|
||||
"",
|
||||
]
|
||||
for page in pages:
|
||||
lines.append(f"## {page['title'] or page['url']}")
|
||||
lines.append("")
|
||||
lines.append(f"Source : {page['url']}")
|
||||
lines.append("")
|
||||
lines.append(page["text"])
|
||||
lines.append("")
|
||||
digest = "\n".join(lines).encode("utf-8")
|
||||
try:
|
||||
saved = save_raw_file(
|
||||
params.vault, params.path, digest, overwrite=True, allow_docs=False
|
||||
)
|
||||
except ServiceError as e:
|
||||
raise ToolError(e.message, code=e.code, details=e.details) from e
|
||||
return {
|
||||
"url": start,
|
||||
"vault": params.vault,
|
||||
"path": saved.get("path", params.path),
|
||||
"pages": len(pages),
|
||||
"failed": failures[:10],
|
||||
"size": saved.get("size", len(digest)),
|
||||
}
|
||||
@@ -0,0 +1,229 @@
|
||||
"""Document-production tools (phase 2 #92) — WRITE, confirmation required.
|
||||
|
||||
The assistant can generate real files inside a vault:
|
||||
|
||||
* ``create_xlsx`` — spreadsheet (openpyxl);
|
||||
* ``create_docx`` — Word document (python-docx);
|
||||
* ``create_csv`` — CSV (stdlib);
|
||||
* ``create_pdf`` — PDF (reportlab, from markdown-ish content).
|
||||
|
||||
Every tool is ``WRITE`` (two-step confirm in the UI / propose-apply over MCP),
|
||||
vault-scoped through ``requires_vault`` and saved via the shared mutation
|
||||
service (path safety, read-only check, backup on overwrite).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import csv as csv_lib
|
||||
import io
|
||||
import logging
|
||||
import re
|
||||
from typing import Any
|
||||
from xml.sax import saxutils
|
||||
|
||||
from backend.services.errors import ServiceError
|
||||
from backend.services.mutations import save_raw_file
|
||||
from backend.tools.context import ToolContext, ToolError, ToolRisk
|
||||
from backend.tools.registry import tool
|
||||
from backend.tools.schemas import CsvInput, DocxInput, PdfInput, SpreadsheetInput
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.documents")
|
||||
|
||||
MAX_PDF_CHARS = 200_000
|
||||
MAX_ROWS = 5_000
|
||||
|
||||
|
||||
def _save(vault: str, path: str, content: bytes, overwrite: bool) -> dict[str, Any]:
|
||||
"""Shared save helper (maps ServiceError to ToolError)."""
|
||||
try:
|
||||
return save_raw_file(vault, path, content, overwrite=overwrite, allow_docs=True)
|
||||
except ServiceError as e:
|
||||
raise ToolError(e.message, code=e.code, details=e.details) from e
|
||||
|
||||
|
||||
def _check_rows(rows: list[list[Any]]) -> None:
|
||||
if not rows:
|
||||
raise ToolError("Aucune ligne fournie", code="invalid_arguments")
|
||||
if len(rows) > MAX_ROWS:
|
||||
raise ToolError(
|
||||
f"Trop de lignes ({len(rows)} > {MAX_ROWS})", code="invalid_arguments"
|
||||
)
|
||||
|
||||
|
||||
def _check_extension(path: str, expected: str) -> str:
|
||||
"""Enforce the document extension; return the normalized path."""
|
||||
path = (path or "").strip()
|
||||
if not path.lower().endswith(expected):
|
||||
raise ToolError(
|
||||
f"Extension attendue : {expected}", code="invalid_arguments"
|
||||
)
|
||||
return path
|
||||
|
||||
|
||||
@tool(
|
||||
name="create_xlsx",
|
||||
description=(
|
||||
"Create an .xlsx spreadsheet in a vault from rows of cell values "
|
||||
"(first row = header). Use for tables, budgets, checklists the user "
|
||||
"asked to turn into an Excel file."
|
||||
),
|
||||
input_model=SpreadsheetInput,
|
||||
risk=ToolRisk.WRITE,
|
||||
requires_vault=True,
|
||||
)
|
||||
def create_xlsx(ctx: ToolContext, params: SpreadsheetInput) -> dict[str, Any]:
|
||||
"""Build the workbook with openpyxl and save it into the vault."""
|
||||
from openpyxl import Workbook
|
||||
|
||||
_check_rows(params.rows)
|
||||
path = _check_extension(params.path, ".xlsx")
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = params.sheet_name[:31] or "Feuille1"
|
||||
for row in params.rows:
|
||||
ws.append(list(row))
|
||||
buffer = io.BytesIO()
|
||||
wb.save(buffer)
|
||||
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
|
||||
|
||||
|
||||
@tool(
|
||||
name="create_docx",
|
||||
description=(
|
||||
"Create a .docx Word document in a vault from an optional title and "
|
||||
"ordered paragraphs. Use for letters, reports, structured drafts."
|
||||
),
|
||||
input_model=DocxInput,
|
||||
risk=ToolRisk.WRITE,
|
||||
requires_vault=True,
|
||||
)
|
||||
def create_docx(ctx: ToolContext, params: DocxInput) -> dict[str, Any]:
|
||||
"""Build the document with python-docx and save it into the vault."""
|
||||
from docx import Document
|
||||
|
||||
if not params.paragraphs:
|
||||
raise ToolError("Aucun paragraphe fourni", code="invalid_arguments")
|
||||
path = _check_extension(params.path, ".docx")
|
||||
doc = Document()
|
||||
if params.title.strip():
|
||||
doc.add_heading(params.title.strip(), level=1)
|
||||
for paragraph in params.paragraphs:
|
||||
doc.add_paragraph(paragraph)
|
||||
buffer = io.BytesIO()
|
||||
doc.save(buffer)
|
||||
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
|
||||
|
||||
|
||||
@tool(
|
||||
name="create_csv",
|
||||
description=(
|
||||
"Create a .csv file in a vault from rows of cell values (first row = "
|
||||
"header). Use for flat data exports, simple tables."
|
||||
),
|
||||
input_model=CsvInput,
|
||||
risk=ToolRisk.WRITE,
|
||||
requires_vault=True,
|
||||
)
|
||||
def create_csv(ctx: ToolContext, params: CsvInput) -> dict[str, Any]:
|
||||
"""Serialize the rows and save the CSV into the vault."""
|
||||
_check_rows(params.rows)
|
||||
path = _check_extension(params.path, ".csv")
|
||||
delimiter = params.delimiter if params.delimiter in (",", ";", "\t") else ","
|
||||
buffer = io.StringIO()
|
||||
writer = csv_lib.writer(buffer, delimiter=delimiter, lineterminator="\n")
|
||||
writer.writerows(params.rows)
|
||||
return _save(params.vault, path, buffer.getvalue().encode("utf-8"), params.overwrite)
|
||||
|
||||
|
||||
_HEADING_RE = re.compile(r"^(#{1,6})\s+(.*)$")
|
||||
|
||||
|
||||
def _markdown_to_flowables(content: str) -> list[tuple[str, str]]:
|
||||
"""Split markdown-ish content into (style, text) blocks for reportlab."""
|
||||
blocks: list[tuple[str, str]] = []
|
||||
for raw_line in content.splitlines():
|
||||
line = raw_line.rstrip()
|
||||
if not line.strip():
|
||||
continue
|
||||
heading = _HEADING_RE.match(line)
|
||||
if heading:
|
||||
blocks.append((f"H{min(3, len(heading.group(1)))}", heading.group(2).strip()))
|
||||
else:
|
||||
blocks.append(("P", line.strip()))
|
||||
return blocks
|
||||
|
||||
|
||||
def _render_markdown_pdf(content: str, title: str) -> bytes | None:
|
||||
"""Render markdown → HTML → PDF through the document-page pipeline.
|
||||
|
||||
Uses the same stack as the « Download PDF » button of the document viewer
|
||||
(mistune with the table plugin + WeasyPrint print CSS), so tables, code
|
||||
blocks and lists are laid out correctly. Returns ``None`` when WeasyPrint
|
||||
is not importable (missing GTK on some hosts) so the caller can fall back
|
||||
to the simplified reportlab renderer.
|
||||
"""
|
||||
try:
|
||||
import mistune
|
||||
|
||||
from backend.pdf_export import build_pdf_html, generate_pdf
|
||||
|
||||
renderer = mistune.create_markdown(
|
||||
escape=False,
|
||||
plugins=["table", "strikethrough", "footnotes", "task_lists"],
|
||||
)
|
||||
html = renderer(content)
|
||||
return generate_pdf(build_pdf_html(html, title), title)
|
||||
except Exception as e:
|
||||
# WeasyPrint loads GTK lazily: a missing native library can surface at
|
||||
# import OR render time. Fall back to the simple renderer either way.
|
||||
logger.warning("WeasyPrint pipeline unavailable for create_pdf: %s", e)
|
||||
return None
|
||||
|
||||
|
||||
def _render_reportlab_pdf(content: str, title: str) -> bytes:
|
||||
"""Fallback renderer (no WeasyPrint): headings + paragraphs, no tables."""
|
||||
from reportlab.lib.pagesizes import A4
|
||||
from reportlab.lib.styles import getSampleStyleSheet
|
||||
from reportlab.platypus import Paragraph, SimpleDocTemplate, Spacer
|
||||
|
||||
styles = getSampleStyleSheet()
|
||||
style_map = {
|
||||
"P": styles["BodyText"],
|
||||
"H1": styles["Heading1"],
|
||||
"H2": styles["Heading2"],
|
||||
"H3": styles["Heading3"],
|
||||
}
|
||||
buffer = io.BytesIO()
|
||||
doc = SimpleDocTemplate(buffer, pagesize=A4, title=title[:200])
|
||||
story: list[Any] = [Paragraph(saxutils.escape(title[:300]), styles["Title"])]
|
||||
for style, line in _markdown_to_flowables(content):
|
||||
story.append(Spacer(1, 4))
|
||||
story.append(Paragraph(saxutils.escape(line), style_map[style]))
|
||||
doc.build(story)
|
||||
return buffer.getvalue()
|
||||
|
||||
|
||||
@tool(
|
||||
name="create_pdf",
|
||||
description=(
|
||||
"Create a .pdf document in a vault from markdown content (headings, "
|
||||
"paragraphs, tables, code blocks, lists). Use for printable "
|
||||
"deliverables; tables are laid out like the document-page PDF export."
|
||||
),
|
||||
input_model=PdfInput,
|
||||
risk=ToolRisk.WRITE,
|
||||
requires_vault=True,
|
||||
)
|
||||
def create_pdf(ctx: ToolContext, params: PdfInput) -> dict[str, Any]:
|
||||
"""Render the content and save the PDF into the vault.
|
||||
|
||||
Primary path: mistune (tables) + WeasyPrint — identical to the viewer's
|
||||
« Download PDF » export. Fallback (WeasyPrint unavailable): simplified
|
||||
reportlab layout without tables.
|
||||
"""
|
||||
path = _check_extension(params.path, ".pdf")
|
||||
content = params.content[:MAX_PDF_CHARS]
|
||||
pdf_bytes = _render_markdown_pdf(content, params.title[:300])
|
||||
if pdf_bytes is None:
|
||||
pdf_bytes = _render_reportlab_pdf(content, params.title[:300])
|
||||
return _save(params.vault, path, pdf_bytes, params.overwrite)
|
||||
@@ -47,6 +47,14 @@ _STEP_LABELS: dict[str, tuple[str, str | None]] = {
|
||||
"restore_backup": ("backup_restore", "path"),
|
||||
"web_search": ("web_search", "query"),
|
||||
"fetch_url": ("fetch_url", "url"),
|
||||
"crawl_site": ("crawl", "url"),
|
||||
"git_list_repos": ("git_repos", "provider"),
|
||||
"git_search_issues": ("git_issues", "query"),
|
||||
"git_get_file": ("git_file", "path"),
|
||||
"create_xlsx": ("xlsx_create", "path"),
|
||||
"create_docx": ("docx_create", "path"),
|
||||
"create_csv": ("csv_create", "path"),
|
||||
"create_pdf": ("pdf_create", "path"),
|
||||
}
|
||||
|
||||
GENERIC_KEY = "generic"
|
||||
|
||||
@@ -251,6 +251,90 @@ class FetchUrlInput(BaseModel):
|
||||
"""Fetch one public web page and return its readable text."""
|
||||
|
||||
url: str = Field(..., description="Absolute http(s) URL of a public page")
|
||||
render: bool = Field(
|
||||
False,
|
||||
description="Render JavaScript with the optional Playwright worker (dynamic SPA pages)",
|
||||
)
|
||||
|
||||
|
||||
class CrawlSiteInput(BaseModel):
|
||||
"""Crawl a small public site (same-host only) and save a digest into a vault."""
|
||||
|
||||
url: str = Field(..., description="Absolute http(s) URL where the crawl starts")
|
||||
vault: str = Field(..., description="Vault name")
|
||||
path: str = Field(..., description="Vault-relative path of the digest file to write (.md)")
|
||||
max_pages: int = Field(5, ge=1, le=20, description="Maximum number of pages to crawl")
|
||||
|
||||
|
||||
class GitProviderInput(BaseModel):
|
||||
"""Base fields for connected-source tools (Gitea / GitHub)."""
|
||||
|
||||
provider: str = Field(..., description="'gitea' (OBSIGATE_GITEA_URL) or 'github'")
|
||||
repo: str = Field("", description="Optional 'owner/name' repository filter")
|
||||
limit: int = Field(20, ge=1, le=50, description="Maximum number of entries")
|
||||
|
||||
|
||||
class GitSearchIssuesInput(BaseModel):
|
||||
"""Search issues/pull requests on a connected Gitea or GitHub instance."""
|
||||
|
||||
provider: str = Field(..., description="'gitea' or 'github'")
|
||||
query: str = Field(..., min_length=1, description="Search keywords")
|
||||
repo: str = Field("", description="Optional 'owner/name' scope (empty = instance-wide)")
|
||||
state: str = Field("open", description="'open' or 'closed'")
|
||||
limit: int = Field(10, ge=1, le=20, description="Maximum number of issues")
|
||||
|
||||
|
||||
class GitGetFileInput(BaseModel):
|
||||
"""Read a file from a connected Gitea or GitHub repository."""
|
||||
|
||||
provider: str = Field(..., description="'gitea' or 'github'")
|
||||
repo: str = Field(..., description="'owner/name' repository")
|
||||
path: str = Field(..., description="Repository-relative file path")
|
||||
ref: str = Field("", description="Optional branch/tag/commit (empty = default branch)")
|
||||
|
||||
|
||||
class SpreadsheetInput(BaseModel):
|
||||
"""Create an .xlsx spreadsheet in a vault from rows of cells."""
|
||||
|
||||
vault: str = Field(..., description="Vault name")
|
||||
path: str = Field(..., description="Vault-relative path of the file to write (.xlsx)")
|
||||
rows: list[list[str | int | float | bool | None]] = Field(
|
||||
..., description="Rows of cell values (first row = header)"
|
||||
)
|
||||
sheet_name: str = Field("Feuille1", description="Worksheet name")
|
||||
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
|
||||
|
||||
|
||||
class DocxInput(BaseModel):
|
||||
"""Create a .docx Word document in a vault from paragraphs."""
|
||||
|
||||
vault: str = Field(..., description="Vault name")
|
||||
path: str = Field(..., description="Vault-relative path of the file to write (.docx)")
|
||||
title: str = Field("", description="Optional document title (heading 1)")
|
||||
paragraphs: list[str] = Field(..., description="Paragraph texts, in order")
|
||||
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
|
||||
|
||||
|
||||
class CsvInput(BaseModel):
|
||||
"""Create a .csv file in a vault from rows of cells."""
|
||||
|
||||
vault: str = Field(..., description="Vault name")
|
||||
path: str = Field(..., description="Vault-relative path of the file to write (.csv)")
|
||||
rows: list[list[str | int | float | bool | None]] = Field(
|
||||
..., description="Rows of cell values (first row = header)"
|
||||
)
|
||||
delimiter: str = Field(",", description="Field separator (',' ';' '\\t')")
|
||||
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
|
||||
|
||||
|
||||
class PdfInput(BaseModel):
|
||||
"""Create a .pdf document in a vault from markdown-ish content."""
|
||||
|
||||
vault: str = Field(..., description="Vault name")
|
||||
path: str = Field(..., description="Vault-relative path of the file to write (.pdf)")
|
||||
title: str = Field("Document", description="Document title")
|
||||
content: str = Field(..., description="Content (headings with #/##, then paragraphs)")
|
||||
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
|
||||
|
||||
|
||||
class ToolResult(BaseModel):
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
"""Tool-layer secrets — user-configured tokens & API keys (#103).
|
||||
|
||||
The connected-source (Gitea / GitHub) and keyed web-search (Tavily, Brave,
|
||||
SerpAPI, Exa) tools read their credentials through this module instead of
|
||||
``os.environ`` directly. The value comes from the store the user edits in the
|
||||
configuration page (``data/api_keys.json`` — the same file the AI provider
|
||||
keys use) first, then falls back to the environment (Infisical-injected in
|
||||
production). Nothing is ever hard-coded and no tool result carries a secret
|
||||
(the registry redacts payloads).
|
||||
|
||||
Allowed names are whitelisted: only the variables below can be stored or
|
||||
deleted from the configuration page.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.secrets")
|
||||
|
||||
# Whitelisted configuration names (config page « Sources connectées & recherche »).
|
||||
TOOL_KEY_NAMES: tuple[str, ...] = (
|
||||
"OBSIGATE_TAVILY_API_KEY",
|
||||
"OBSIGATE_BRAVE_API_KEY",
|
||||
"OBSIGATE_SERPAPI_API_KEY",
|
||||
"OBSIGATE_EXA_API_KEY",
|
||||
"OBSIGATE_GITEA_URL",
|
||||
"OBSIGATE_GITEA_TOKEN",
|
||||
"OBSIGATE_GITHUB_TOKEN",
|
||||
)
|
||||
|
||||
_SECRET_MARKERS = ("API_KEY", "TOKEN")
|
||||
|
||||
|
||||
def _keys_file() -> Path:
|
||||
base = os.environ.get("OBSIGATE_DATA_DIR", "data")
|
||||
return Path(base) / "api_keys.json"
|
||||
|
||||
|
||||
def _read_keys() -> dict:
|
||||
path = _keys_file()
|
||||
if not path.exists():
|
||||
return {}
|
||||
try:
|
||||
data = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError) as e:
|
||||
logger.warning("tool key store unreadable (%s): %s", path, e)
|
||||
return {}
|
||||
return data if isinstance(data, dict) else {}
|
||||
|
||||
|
||||
def _write_keys(data: dict) -> None:
|
||||
path = _keys_file()
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = path.with_suffix(".tmp")
|
||||
tmp.write_text(json.dumps(data, indent=2), encoding="utf-8")
|
||||
tmp.replace(path)
|
||||
|
||||
|
||||
def is_secret_name(name: str) -> bool:
|
||||
"""True for API keys / tokens (masked in API responses); URLs are clear."""
|
||||
return any(marker in name for marker in _SECRET_MARKERS)
|
||||
|
||||
|
||||
def mask_value(name: str, value: str) -> str:
|
||||
"""Mask a secret for display; non-secret values (URLs) are returned as-is."""
|
||||
if not value:
|
||||
return ""
|
||||
if not is_secret_name(name):
|
||||
return value
|
||||
return value[:4] + "..." + value[-4:] if len(value) > 8 else "***"
|
||||
|
||||
|
||||
def get_tool_key(name: str) -> str:
|
||||
"""Stored (configuration page) value first, then environment fallback."""
|
||||
if name not in TOOL_KEY_NAMES:
|
||||
return os.environ.get(name, "").strip()
|
||||
stored = _read_keys().get(name)
|
||||
if isinstance(stored, str) and stored.strip():
|
||||
return stored.strip()
|
||||
return os.environ.get(name, "").strip()
|
||||
|
||||
|
||||
def set_tool_key(name: str, value: str) -> None:
|
||||
"""Persist one whitelisted key into the store (admin configuration page)."""
|
||||
if name not in TOOL_KEY_NAMES:
|
||||
raise ValueError(f"Clé non prise en charge: {name}")
|
||||
value = (value or "").strip()
|
||||
keys = _read_keys()
|
||||
if value:
|
||||
keys[name] = value
|
||||
else:
|
||||
keys.pop(name, None)
|
||||
_write_keys(keys)
|
||||
|
||||
|
||||
def delete_tool_key(name: str) -> bool:
|
||||
"""Remove one key from the store; return True when it existed."""
|
||||
if name not in TOOL_KEY_NAMES:
|
||||
raise ValueError(f"Clé non prise en charge: {name}")
|
||||
keys = _read_keys()
|
||||
if name in keys:
|
||||
del keys[name]
|
||||
_write_keys(keys)
|
||||
return True
|
||||
return False
|
||||
+181
-4
@@ -8,6 +8,16 @@ Phase 1 of the documented web-toolset roadmap:
|
||||
answering « je n'ai pas accès à internet ».
|
||||
* ``fetch_url`` — retrieve a public web page and return readable text.
|
||||
|
||||
Phase 2 (#92) additions:
|
||||
|
||||
* keyed providers — Tavily, Brave Search, SerpAPI and Exa are used first when
|
||||
their API key is configured (env, injected by Infisical in production);
|
||||
* SQLite cache — search/fetch results are cached with a TTL
|
||||
(:mod:`backend.tools.webcache`);
|
||||
* retry with backoff — transient network errors get one extra attempt;
|
||||
* dynamic rendering — ``fetch_url(render=True)`` uses an isolated Playwright
|
||||
worker (optional dependency, graceful degradation).
|
||||
|
||||
All are READ-risk tools (no confirmation), rate-limited through the shared
|
||||
registry, SSRF-guarded (scheme + private-address rejection), and size-capped.
|
||||
|
||||
@@ -16,6 +26,14 @@ Configuration (environment):
|
||||
* ``OBSIGATE_WEB_TIMEOUT`` — seconds, default 10
|
||||
* ``OBSIGATE_WEB_FALLBACK`` — ``0``/``false`` disables the keyless HTML
|
||||
fallbacks (SearXNG only), default enabled
|
||||
* ``OBSIGATE_TAVILY_API_KEY`` / ``OBSIGATE_BRAVE_API_KEY`` /
|
||||
``OBSIGATE_SERPAPI_API_KEY`` / ``OBSIGATE_EXA_API_KEY`` — optional keyed
|
||||
providers, tried before SearXNG when set
|
||||
* ``OBSIGATE_WEB_PROVIDERS`` — optional comma-separated provider order
|
||||
(e.g. ``brave,searxng``); keyed providers without a key are skipped
|
||||
* ``OBSIGATE_WEB_RETRY`` — extra attempts for transient network errors
|
||||
(default 1)
|
||||
* ``OBSIGATE_WEB_CACHE_TTL`` — cache TTL seconds, ``0`` disables (default 900)
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -28,15 +46,18 @@ import logging
|
||||
import os
|
||||
import re
|
||||
import socket
|
||||
import time
|
||||
from collections.abc import Callable
|
||||
from typing import Any
|
||||
from urllib.parse import parse_qs, urlparse
|
||||
|
||||
import httpx
|
||||
|
||||
from backend.tools import webcache
|
||||
from backend.tools.context import ToolError, ToolRisk, ToolScope
|
||||
from backend.tools.registry import tool
|
||||
from backend.tools.schemas import FetchUrlInput, WebSearchInput
|
||||
from backend.tools.secrets import get_tool_key
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.web")
|
||||
|
||||
@@ -48,6 +69,7 @@ WEB_FALLBACK_ENABLED = os.environ.get("OBSIGATE_WEB_FALLBACK", "1").strip().lowe
|
||||
"no",
|
||||
"off",
|
||||
}
|
||||
WEB_RETRY_ATTEMPTS = int(os.environ.get("OBSIGATE_WEB_RETRY", "1"))
|
||||
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
|
||||
# Search engines reject non-browser agents on their public HTML endpoints.
|
||||
BROWSER_UA = (
|
||||
@@ -163,6 +185,115 @@ def _result(
|
||||
}
|
||||
|
||||
|
||||
def _with_retry(call: Callable[[], Any]) -> Any:
|
||||
"""Run *call* with one extra attempt on transient network errors.
|
||||
|
||||
House-made backoff (the roadmap's « tenacity ou boucle maison »): DNS
|
||||
blips and rate-limit hiccups are the common failure mode, and a single
|
||||
retry keeps the fallback chain from being consumed too early.
|
||||
"""
|
||||
for attempt in range(1 + max(0, WEB_RETRY_ATTEMPTS)):
|
||||
try:
|
||||
return call()
|
||||
except httpx.TransportError:
|
||||
if attempt >= max(0, WEB_RETRY_ATTEMPTS):
|
||||
raise
|
||||
time.sleep(0.2 * (attempt + 1))
|
||||
raise RuntimeError("unreachable") # pragma: no cover
|
||||
|
||||
|
||||
def _env_key(name: str) -> str:
|
||||
"""Read an API key: configuration-page store first, then environment."""
|
||||
return get_tool_key(name)
|
||||
|
||||
|
||||
def _search_tavily(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
"""Tavily Search API (agent-oriented results, key required)."""
|
||||
resp = httpx.post(
|
||||
"https://api.tavily.com/search",
|
||||
json={
|
||||
"api_key": _env_key("OBSIGATE_TAVILY_API_KEY"),
|
||||
"query": query,
|
||||
"max_results": params.max_results,
|
||||
"search_depth": "basic",
|
||||
"include_answer": False,
|
||||
},
|
||||
headers={"User-Agent": USER_AGENT},
|
||||
timeout=WEB_TIMEOUT,
|
||||
)
|
||||
resp.raise_for_status()
|
||||
data = resp.json()
|
||||
return [
|
||||
_result(item.get("title") or "", item.get("url") or "", item.get("content") or "")
|
||||
for item in (data.get("results") or [])
|
||||
], []
|
||||
|
||||
|
||||
def _search_brave(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
"""Brave Search API (key required)."""
|
||||
resp = httpx.get(
|
||||
"https://api.search.brave.com/res/v1/web/search",
|
||||
params={"q": query, "count": params.max_results, "safesearch": "moderate"},
|
||||
headers={
|
||||
"X-Subscription-Id": _env_key("OBSIGATE_BRAVE_API_KEY"),
|
||||
"Accept": "application/json",
|
||||
"User-Agent": USER_AGENT,
|
||||
},
|
||||
timeout=WEB_TIMEOUT,
|
||||
)
|
||||
resp.raise_for_status()
|
||||
data = resp.json()
|
||||
return [
|
||||
_result(item.get("title") or "", item.get("url") or "", item.get("description") or "")
|
||||
for item in ((data.get("web") or {}).get("results") or [])
|
||||
], []
|
||||
|
||||
|
||||
def _search_serpapi(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
"""SerpAPI (Google SERP, key required)."""
|
||||
resp = httpx.get(
|
||||
"https://serpapi.com/search",
|
||||
params={"q": query, "api_key": _env_key("OBSIGATE_SERPAPI_API_KEY"),
|
||||
"num": params.max_results},
|
||||
headers={"User-Agent": USER_AGENT},
|
||||
timeout=WEB_TIMEOUT,
|
||||
)
|
||||
resp.raise_for_status()
|
||||
data = resp.json()
|
||||
return [
|
||||
_result(item.get("title") or "", item.get("link") or "", item.get("snippet") or "")
|
||||
for item in (data.get("organic_results") or [])
|
||||
], []
|
||||
|
||||
|
||||
def _search_exa(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
"""Exa neural search (key required)."""
|
||||
resp = httpx.post(
|
||||
"https://api.exa.ai/search",
|
||||
json={"query": query, "numResults": params.max_results},
|
||||
headers={
|
||||
"x-api-key": _env_key("OBSIGATE_EXA_API_KEY"),
|
||||
"User-Agent": USER_AGENT,
|
||||
},
|
||||
timeout=WEB_TIMEOUT,
|
||||
)
|
||||
resp.raise_for_status()
|
||||
data = resp.json()
|
||||
return [
|
||||
_result(item.get("title") or "", item.get("url") or "", (item.get("text") or "")[:600])
|
||||
for item in (data.get("results") or [])
|
||||
], []
|
||||
|
||||
|
||||
# Keyed providers: name -> (implementation, API key env var)
|
||||
_KEYED_PROVIDERS: dict[str, tuple[_Provider, str]] = {
|
||||
"tavily": (_search_tavily, "OBSIGATE_TAVILY_API_KEY"),
|
||||
"brave": (_search_brave, "OBSIGATE_BRAVE_API_KEY"),
|
||||
"serpapi": (_search_serpapi, "OBSIGATE_SERPAPI_API_KEY"),
|
||||
"exa": (_search_exa, "OBSIGATE_EXA_API_KEY"),
|
||||
}
|
||||
|
||||
|
||||
def _search_searxng(
|
||||
query: str, params: WebSearchInput
|
||||
) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
@@ -288,8 +419,22 @@ _Provider = Callable[[str, WebSearchInput], "tuple[list[dict[str, Any]], list[st
|
||||
|
||||
|
||||
def _provider_chain() -> list[tuple[str, _Provider]]:
|
||||
"""Ordered providers: self-hosted meta-search first, then keyless fallbacks."""
|
||||
chain: list[tuple[str, _Provider]] = [("searxng", _search_searxng)]
|
||||
"""Ordered providers: keyed APIs first, then self-hosted, then keyless.
|
||||
|
||||
``OBSIGATE_WEB_PROVIDERS`` (comma-separated) overrides the default order;
|
||||
unknown names are ignored and keyed providers without their key are skipped.
|
||||
"""
|
||||
chain: list[tuple[str, _Provider]] = []
|
||||
configured = [
|
||||
name.strip().lower()
|
||||
for name in os.environ.get("OBSIGATE_WEB_PROVIDERS", "").split(",")
|
||||
if name.strip()
|
||||
]
|
||||
for name in configured or list(_KEYED_PROVIDERS):
|
||||
entry = _KEYED_PROVIDERS.get(name)
|
||||
if entry and _env_key(entry[1]):
|
||||
chain.append((name, entry[0]))
|
||||
chain.append(("searxng", _search_searxng))
|
||||
if WEB_FALLBACK_ENABLED:
|
||||
chain.append(("duckduckgo", _search_duckduckgo))
|
||||
chain.append(("bing", _search_bing))
|
||||
@@ -313,6 +458,17 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
|
||||
if not query:
|
||||
raise ToolError("Requête vide", code="invalid_arguments")
|
||||
|
||||
key = webcache.cache_key("search", {
|
||||
"q": query,
|
||||
"max_results": params.max_results,
|
||||
"category": params.category,
|
||||
"language": params.language,
|
||||
"page": params.page,
|
||||
})
|
||||
cached = webcache.cache_get(key)
|
||||
if cached is not None:
|
||||
return {**cached, "cached": True}
|
||||
|
||||
attempts: list[str] = []
|
||||
unresponsive: list[str] = []
|
||||
reachable = False
|
||||
@@ -320,8 +476,12 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
|
||||
|
||||
for name, provider in _provider_chain():
|
||||
attempts.append(name)
|
||||
|
||||
def _attempt(p: _Provider = provider) -> tuple[list[dict[str, Any]], list[str]]:
|
||||
return p(query, params)
|
||||
|
||||
try:
|
||||
results, engines = provider(query, params)
|
||||
results, engines = _with_retry(_attempt)
|
||||
except (httpx.HTTPError, ValueError, AttributeError) as e:
|
||||
logger.warning("web_search provider %s failed: %s", name, e)
|
||||
last_error = e
|
||||
@@ -338,6 +498,7 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
|
||||
}
|
||||
if unresponsive:
|
||||
payload["unresponsive_engines"] = unresponsive[:8]
|
||||
webcache.cache_set(key, payload)
|
||||
return payload
|
||||
|
||||
if not reachable:
|
||||
@@ -378,6 +539,20 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
|
||||
def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
|
||||
"""Retrieve one page, guard against SSRF, and extract its text."""
|
||||
url = _assert_public_http_url(params.url.strip())
|
||||
key = webcache.cache_key("fetch", {"url": url, "render": params.render})
|
||||
cached = webcache.cache_get(key)
|
||||
if cached is not None:
|
||||
return {**cached, "cached": True}
|
||||
|
||||
if params.render:
|
||||
# Dynamic pages (SPA/React): delegated to the isolated Playwright
|
||||
# worker; the browser dependency stays optional (graceful error).
|
||||
from backend.tools.webrender import render_page
|
||||
|
||||
payload = render_page(url)
|
||||
webcache.cache_set(key, payload)
|
||||
return payload
|
||||
|
||||
try:
|
||||
# Follow redirects manually so every hop is re-checked against the
|
||||
# private-address SSRF guard (a public page can redirect to 127.0.0.1).
|
||||
@@ -414,10 +589,12 @@ def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
|
||||
title_match = re.search(r"<title[^>]*>(.*?)</title>", raw, re.IGNORECASE | re.DOTALL)
|
||||
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
|
||||
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
|
||||
return {
|
||||
payload = {
|
||||
"url": str(resp.url),
|
||||
"status": resp.status_code,
|
||||
"title": title,
|
||||
"text": text,
|
||||
"truncated": len(raw) > MAX_TEXT_CHARS,
|
||||
}
|
||||
webcache.cache_set(key, payload)
|
||||
return payload
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
"""SQLite cache for web tool results (search results, fetched pages).
|
||||
|
||||
Phase 2 of the web-toolset roadmap (« Transverse »): repeated web searches and
|
||||
page fetches (common in agent loops, where the model re-reads a source) must
|
||||
not hammer the providers. Results are cached in a dedicated SQLite table with
|
||||
a TTL; the cache is best-effort — any error silently disables it so a broken
|
||||
database file never takes the assistant down.
|
||||
|
||||
Configuration (environment):
|
||||
* ``OBSIGATE_DATA_DIR`` — base data directory (default ``data``)
|
||||
* ``OBSIGATE_WEB_CACHE_PATH`` — explicit cache file override
|
||||
* ``OBSIGATE_WEB_CACHE_TTL`` — seconds, ``0`` disables the cache (default 900)
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import sqlite3
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.webcache")
|
||||
|
||||
DEFAULT_TTL_SECONDS = 900
|
||||
_schema_ready = False
|
||||
_write_lock = threading.Lock()
|
||||
|
||||
|
||||
def ttl_seconds() -> float:
|
||||
"""Configured TTL in seconds (``0`` = cache disabled)."""
|
||||
return float(os.environ.get("OBSIGATE_WEB_CACHE_TTL", str(DEFAULT_TTL_SECONDS)))
|
||||
|
||||
|
||||
def _cache_path() -> Path:
|
||||
override = os.environ.get("OBSIGATE_WEB_CACHE_PATH", "").strip()
|
||||
if override:
|
||||
return Path(override)
|
||||
return Path(os.environ.get("OBSIGATE_DATA_DIR", "data")) / "web_cache.sqlite3"
|
||||
|
||||
|
||||
def _connect() -> sqlite3.Connection:
|
||||
"""Open (and lazily create) the cache database."""
|
||||
global _schema_ready
|
||||
path = _cache_path()
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
conn = sqlite3.connect(path, timeout=5, check_same_thread=False)
|
||||
if not _schema_ready:
|
||||
conn.execute(
|
||||
"CREATE TABLE IF NOT EXISTS web_cache ("
|
||||
"key TEXT PRIMARY KEY, value TEXT NOT NULL, created REAL NOT NULL)"
|
||||
)
|
||||
conn.commit()
|
||||
_schema_ready = True
|
||||
return conn
|
||||
|
||||
|
||||
def cache_key(prefix: str, payload: dict[str, Any]) -> str:
|
||||
"""Deterministic cache key from a prefix and the normalized arguments."""
|
||||
raw = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str)
|
||||
digest = hashlib.sha256(raw.encode("utf-8")).hexdigest()[:32]
|
||||
return f"{prefix}:{digest}"
|
||||
|
||||
|
||||
def cache_get(key: str) -> Any | None:
|
||||
"""Return the cached payload for *key*, or ``None`` (miss/expiry/disabled)."""
|
||||
if ttl_seconds() <= 0:
|
||||
return None
|
||||
try:
|
||||
conn = _connect()
|
||||
row = conn.execute(
|
||||
"SELECT value, created FROM web_cache WHERE key = ?", (key,)
|
||||
).fetchone()
|
||||
conn.close()
|
||||
except sqlite3.Error as e:
|
||||
logger.warning("web cache read failed (%s): %s", key, e)
|
||||
return None
|
||||
if row is None:
|
||||
return None
|
||||
value, created = row
|
||||
if time.time() - float(created) > ttl_seconds():
|
||||
return None
|
||||
try:
|
||||
return json.loads(value)
|
||||
except (ValueError, TypeError):
|
||||
return None
|
||||
|
||||
|
||||
def cache_set(key: str, value: Any) -> None:
|
||||
"""Store *value* under *key* (best effort, never raises)."""
|
||||
if ttl_seconds() <= 0:
|
||||
return
|
||||
try:
|
||||
with _write_lock:
|
||||
conn = _connect()
|
||||
conn.execute(
|
||||
"INSERT INTO web_cache (key, value, created) VALUES (?, ?, ?) "
|
||||
"ON CONFLICT(key) DO UPDATE SET value = excluded.value, created = excluded.created",
|
||||
(key, json.dumps(value, ensure_ascii=False, default=str), time.time()),
|
||||
)
|
||||
conn.commit()
|
||||
conn.close()
|
||||
except sqlite3.Error as e:
|
||||
logger.warning("web cache write failed (%s): %s", key, e)
|
||||
|
||||
|
||||
def purge_expired() -> int:
|
||||
"""Delete expired rows; return the number of removed entries (maintenance)."""
|
||||
try:
|
||||
conn = _connect()
|
||||
cursor = conn.execute(
|
||||
"DELETE FROM web_cache WHERE created < ?", (time.time() - ttl_seconds(),)
|
||||
)
|
||||
conn.commit()
|
||||
deleted = cursor.rowcount
|
||||
conn.close()
|
||||
return int(deleted)
|
||||
except sqlite3.Error as e:
|
||||
logger.warning("web cache purge failed: %s", e)
|
||||
return 0
|
||||
|
||||
|
||||
def clear_cache() -> int:
|
||||
"""Drop every cached entry (tests / admin); returns the number of rows."""
|
||||
try:
|
||||
conn = _connect()
|
||||
cursor = conn.execute("DELETE FROM web_cache")
|
||||
conn.commit()
|
||||
deleted = cursor.rowcount
|
||||
conn.close()
|
||||
return int(deleted)
|
||||
except sqlite3.Error as e:
|
||||
logger.warning("web cache clear failed: %s", e)
|
||||
return 0
|
||||
@@ -0,0 +1,100 @@
|
||||
"""Dynamic page rendering (Playwright) — ``fetch_url(render=True)``.
|
||||
|
||||
Static pages are fetched with httpx inside :mod:`backend.tools.web`. Dynamic
|
||||
pages (SPA/React, JS-loaded content) need a real browser engine; this module
|
||||
runs one Playwright call inside a dedicated worker thread so browser
|
||||
crashes/timeouts never take over the tool layer, and the heavyweight
|
||||
dependency stays optional:
|
||||
|
||||
* not installed → ``ToolError(code="playwright_unavailable")`` with a clear
|
||||
message (the assistant explains the limitation instead of hanging);
|
||||
* installed → ``pip install playwright && playwright install chromium``.
|
||||
|
||||
The SSRF guard (scheme + private-address rejection) is applied before the
|
||||
browser navigates. Note: unlike the httpx path, internal redirects performed
|
||||
by the browser engine are not re-checked hop by hop.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import html as html_lib
|
||||
import logging
|
||||
import re
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from typing import Any
|
||||
|
||||
from backend.tools.context import ToolError
|
||||
from backend.tools.web import (
|
||||
MAX_TEXT_CHARS,
|
||||
USER_AGENT,
|
||||
_assert_public_http_url,
|
||||
_html_to_text,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("obsigate.tools.webrender")
|
||||
|
||||
# One worker: browser automation is serialized on purpose (one Chromium at a
|
||||
# time keeps memory predictable on small hosts).
|
||||
_executor = ThreadPoolExecutor(max_workers=1, thread_name_prefix="obsigate-playwright")
|
||||
GOTO_TIMEOUT_MS = 20_000
|
||||
|
||||
|
||||
def _playwright_available() -> bool:
|
||||
try:
|
||||
import playwright # noqa: F401
|
||||
except ImportError:
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
def _render_in_worker(url: str) -> dict[str, Any]:
|
||||
"""Synchronous Playwright render — runs in the dedicated worker thread."""
|
||||
from playwright.sync_api import sync_playwright
|
||||
|
||||
status = 0
|
||||
with sync_playwright() as p:
|
||||
browser = p.chromium.launch(headless=True)
|
||||
try:
|
||||
page = browser.new_page(user_agent=USER_AGENT)
|
||||
response = page.goto(url, wait_until="networkidle", timeout=GOTO_TIMEOUT_MS)
|
||||
if response is not None:
|
||||
status = response.status
|
||||
raw = page.content()
|
||||
title = html_lib.unescape(page.title() or "").strip()
|
||||
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
|
||||
finally:
|
||||
browser.close()
|
||||
title = re.sub(r"\s+", " ", title)[:300]
|
||||
return {
|
||||
"url": url,
|
||||
"status": status,
|
||||
"title": title,
|
||||
"text": text,
|
||||
"rendered": True,
|
||||
"truncated": len(raw) > MAX_TEXT_CHARS,
|
||||
}
|
||||
|
||||
|
||||
def render_page(url: str) -> dict[str, Any]:
|
||||
"""Render *url* (JavaScript included) and return readable text.
|
||||
|
||||
Raises:
|
||||
ToolError: ``playwright_unavailable`` when the optional dependency is
|
||||
missing, ``render_unavailable`` when the render itself failed.
|
||||
"""
|
||||
_assert_public_http_url(url)
|
||||
if not _playwright_available():
|
||||
raise ToolError(
|
||||
"Rendu dynamique indisponible : Playwright n'est pas installé "
|
||||
"(pip install playwright && playwright install chromium).",
|
||||
code="playwright_unavailable",
|
||||
)
|
||||
try:
|
||||
return _executor.submit(_render_in_worker, url).result(timeout=GOTO_TIMEOUT_MS / 1000 + 40)
|
||||
except ToolError:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.warning("render_page failed for %s: %s", url, e)
|
||||
raise ToolError(
|
||||
"Le rendu dynamique de la page a échoué.", code="render_unavailable"
|
||||
) from e
|
||||
Generated
+1
-1
@@ -2626,7 +2626,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "obsigate-desktop"
|
||||
version = "2.9.0"
|
||||
version = "2.11.5"
|
||||
dependencies = [
|
||||
"chrono",
|
||||
"env_logger",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "obsigate-desktop"
|
||||
version = "2.9.0"
|
||||
version = "2.11.5"
|
||||
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
|
||||
authors = ["Bruno Charest"]
|
||||
edition = "2021"
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
|
||||
"productName": "ObsiGate",
|
||||
"version": "2.9.0",
|
||||
"version": "2.11.5",
|
||||
"identifier": "com.obsigate.desktop",
|
||||
"build": {
|
||||
"frontendDist": "../frontend",
|
||||
|
||||
+17
-6
@@ -144,12 +144,12 @@ Avant de corriger quoi que ce soit, un agent IA doit :
|
||||
| *BUG-032* | [🟡 IMPORTANT] Indexation : symlinks suivis (contenu hors vault indexé) + scan initial coûteux | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/indexer.py` | Placer un symlink dans le vault vers un dossier externe puis relancer l'index | `_scan_vault` réécrit avec `os.walk(followlinks=False)` + refus des symlinks sortant de la racine ; test `TestSymlinkIndexing` | Scan incrémental/index persistant : voir #86 (phase 3) |
|
||||
| *BUG-033* | [🟡 IMPORTANT] Recherche classique et tool IA `search_fulltext` en O(N) sans inverted index | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/search.py`, `backend/tools/service.py` | `GET /api/search` sur un vault de 50 000 fichiers | `search()` récupère les candidats via l'inverted index (intersection des termes + expansion de préfixes), repli sur le scan pendant la construction | `search_fulltext` en bénéficie automatiquement |
|
||||
| *BUG-034* | [🟡 IMPORTANT] CSP affaiblie (`'unsafe-inline'` + CDN distants) et token d'accès en sessionStorage | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/main.py`, `frontend/js/auth.js`, `frontend/js/admin.js`, `frontend/js/sync.js` | Inspecter les en-têtes CSP ; lire sessionStorage en console | Token en mémoire + cookie HttpOnly (plus de `sessionStorage`) ; CSP durcie (`object-src 'none'`, `base-uri`, `form-action`, `frame-ancestors`). *Reste : migration nonce* | `'unsafe-inline'` conservé tant que les gestionnaires inline n'ont pas été convertis (résidu documenté) |
|
||||
| *BUG-035* | [🔵 MINEUR] `secret_redactor` : faux positifs sur les hashs hex (git, SHA) | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/secret_redactor.py` | Lire une note contenant un commit git (40 caractères hexadécimaux) | Restreindre le périmètre de détection (contexte clé/token) + whitelist | Contenus mutilés dans les lectures et réponses IA |
|
||||
| *BUG-036* | [🔵 MINEUR] Collab WebSocket : token en query string | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/collab.py` | Observer l'URL du websocket dans le trafic réseau | Passer le token en header / étape d'authentification initiale ; borner la taille des messages | Jeton visible dans les logs/proxys |
|
||||
| *BUG-037* | [🔵 MINEUR] Compte « anonymous » administrateur si auth désactivée | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/middleware.py` | Démarrer avec l'authentification désactivée | Avertissement explicite au démarrage + refus de déploiement public sans auth | Comportement par conception mais risqué si mal configuré |
|
||||
| *BUG-038* | [🔵 MINEUR] Argon2 à 64 MB par vérification : risque d'épuisement mémoire | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/password.py:8` | Lancer de nombreux `POST /api/auth/login` simultanés | Recalibrer (~19 MB, t=2, p=1, norme OWASP actuelle) + maintien du rate-limit | DoS mémoire possible sur les petites instances |
|
||||
| *BUG-039* | [🔵 MINEUR] Enumération de comptes : 429 (verrouillé) vs 401 (inconnu) | 🔴 ouvert | P3 | 🔐 sécurité | IA | `backend/auth/router.py:120` | Tenter un login sur un compte verrouillé puis un nom inconnu | Répondre 401 uniforme avec un timing équivalent | Le statut HTTP distingue l'existence d'un compte |
|
||||
| *BUG-040* | [🔵 MINEUR] Extraction PDF intégrale (100 ko) au scan de démarrage | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/indexer.py:465` | Démarrer sur un vault contenant de nombreux PDF | Analyser les PDF en tâche de fond / à la demande (lazy) | Ralentit fortement le démarrage et le rebuild d'index |
|
||||
| *BUG-035* | [🔵 MINEUR] `secret_redactor` : faux positifs sur les hashs hex (git, SHA) | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/secret_redactor.py` | Lire une note contenant un commit git (40 caractères hexadécimaux) | Masquage hex conditionné au contexte (`_redact_bare_hex_secrets`) : secret exigé dans les 60 caractères précédents, exemption explicite pour `commit`/`sha*`/`hash`/`checksum`/`git`/`etag`. Tests : `tests/test_api_main.py::TestSecretRedactor` (+4) | Contenus mutilés dans les lectures et réponses IA |
|
||||
| *BUG-036* | [🔵 MINEUR] Collab WebSocket : token en query string | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/collab.py` | Observer l'URL du websocket dans le trafic réseau | `authenticate_websocket` ne lit plus `?token=` : cookie HttpOnly `access_token` uniquement ; rejet des trames > `MAX_MESSAGE_CHARS` (16 Mio) avant analyse. Tests : `tests/test_collab.py` (+3) | Jeton visible dans les logs/proxys |
|
||||
| *BUG-037* | [🔵 MINEUR] Compte « anonymous » administrateur si auth désactivée | 🟢 corrigé | P2 | 🔐 sécurité | IA | `backend/auth/middleware.py`, `backend/main.py` | Démarrer avec l'authentification désactivée | `_guard_insecure_auth()` : avertissement explicite + refus de démarrage sur bind non-loopback sans `OBSIGATE_ALLOW_INSECURE=true`. Tests : `tests/test_auth.py::TestInsecureAuthGuard` (+6) | Comportement par conception mais risqué si mal configuré |
|
||||
| *BUG-038* | [🔵 MINEUR] Argon2 à 64 MB par vérification : risque d'épuisement mémoire | 🟢 corrigé | P2 | 🔐 sécurité | IA | `backend/auth/password.py` | Lancer de nombreux `POST /api/auth/login` simultanés | Recalibré à `m=19456 Kio (19 Mio), t=2, p=1` (OWASP) ; anciens hachages valides + rehash auto. Test : `tests/test_auth.py::TestPasswordHashing::test_argon2_memory_recalibrated` | DoS mémoire possible sur les petites instances |
|
||||
| *BUG-039* | [🔵 MINEUR] Enumération de comptes : 429 (verrouillé) vs 401 (inconnu) | 🟢 corrigé | P3 | 🔐 sécurité | IA | `backend/auth/router.py` | Tenter un login sur un compte verrouillé puis un nom inconnu | Login uniforme : inconnu / désactivé / verrouillé / rate-limit par compte → `401 Identifiants invalides` + hachage factice (timing équivalent) ; seul le rate-limit IP reste `429`. Tests : `tests/test_auth_api.py` (+3) | Le statut HTTP distinguait l'existence d'un compte |
|
||||
| *BUG-040* | [🔵 MINEUR] Extraction PDF intégrale (100 ko) au scan de démarrage | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/indexer.py`, `backend/main.py` | Démarrer sur un vault contenant de nombreux PDF | `_scan_vault` ne lit que les métadonnées ; `enrich_pdf_texts()` extrait le texte après l'index (démarrage) et après chaque réindexation. Tests : `tests/test_pdf.py` (+3) | Ralentit fortement le démarrage et le rebuild d'index |
|
||||
| *BUG-041* | [🟡 IMPORTANT] Assistant IA : échec sur un répertoire vide (« Aucun fichier markdown trouvé dans ce dossier ») au lieu de répondre | 🟢 corrigé | P1 | 📱 frontend + ⚙️ backend | IA | `backend/bookslm_routes.py`, `backend/bookslm.py`, `frontend/js/bookslm.js` | Ouvrir l'assistant sur un dossier vide puis envoyer une question | `_resolve_system_prompt` dégrade vers le prompt Général + bloc « Dossier vide » (plus de 404) ; contexte applicatif `app_context` enrichi (documents ouverts, répertoire, recherche, fichiers récents) | Le 404 bloquait toute la requête. Feature #88, fiche `docs/features/ai-app-context.md`. Tests : `tests/test_bookslm.py` (+3), `tests/frontend/ai.test.mjs` |
|
||||
| *BUG-042* | [🟡 IMPORTANT] Assistant IA : liens de fichiers non fiables (« File not found: ») — pas de règle déterministe nom / dossier / chemin | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js` | Cliquer les liens de fichiers/dossiers dans une réponse de l'assistant (noms avec espaces et/ou accents, chemin préfixé par le nom du vault) | `_classifyPath` distingue `name` (copie presse-papiers) / `dir` (révélation arborescence) / `file` (ouverture) ; `_activatePath()` résout le chemin contre l'index du vault (exact → suffixe → basename unique) avant d'agir ; espaces + accents pris en charge (classes Unicode `\p{L}\p{N}\p{M}`, comparaison normalisée NFC, markdown `<…>`/`%20`, code inline, mentions brutes confirmées par l'index) ; `_splitVaultPrefix` retire un préfixe `Vault/…` et ouvre dans ce vault (`_fetchPathsForVault`) | Les liens morts ouvraient un fichier inexistant. Feature #88. Tests : `tests/frontend/ai.test.mjs` (+11) |
|
||||
| *BUG-043* | [🟡 IMPORTANT] Assistant IA : la liste des fournisseurs de la barre latérale ne suit pas les ajouts/retraits de clés API dans la configuration du projet | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/ai.js`, `frontend/js/bookslm.js`, `frontend/js/config.js` | Ajouter (ou supprimer) une clé de fournisseur AI dans la configuration puis observer le menu Fournisseur de l'assistant sans recharger la page | Le picker lit `/api/ai/status` **une seule fois**, à sa construction, et le panneau de l'assistant est un singleton monté pour toute la session → liste figée. Nouveau `refreshAIPickers()` (exporté par `ai.js`) qui reconstruit chaque picker monté dans son emplacement `.ai-picker-slot` (conservé même sans fournisseur configuré, donc un premier fournisseur s'y monte aussi) ; appelé après `saveAIKeys()` et `deleteAIKey()` (`config.js`) ; une sélection dont le fournisseur n'est plus configuré est purgée de `obsigate_ai_picker` (retour au défaut + modèle effacé au lieu d'un nom fantôme) | Il fallait recharger la page pour voir un nouveau fournisseur (ou en voir disparaître un). Feature #82. Tests : `tests/frontend/ai.test.mjs` (+4) |
|
||||
@@ -167,6 +167,11 @@ Avant de corriger quoi que ce soit, un agent IA doit :
|
||||
| *BUG-055* | [🟡 IMPORTANT] Éditeur Forge : l'autocomplétion (Tab) ajoute des espaces parasites, l'effacement détruit le mot complété et la complétion fantôme est illisible | 🟢 corrigé | P1 | 📱 frontend + ⚙️ backend | IA | `frontend/editor-poc.html`, `frontend/js/autocomplete.js`, `backend/ai.py`, `.gitea/workflows/ci.yml`, `tests/frontend/forge-completion.test.mjs` (nouveau) | Forge : taper un mot, puis Tab pour compléter ; un espace (voire deux) s'insère avant le mot complété, et le retour arrière efface l'ajout. La prédiction IA s'affichait décalée (texte miroir du document entier) | **Cause** : trois gestionnaires `keydown` Tab indépendants s'exécutaient tous — l'indentation (`insertAtCursor(' ')`) s'ajoutait à la complétion de mot et à l'acceptation du ghost. **Correctif** : gestion **unifiée** de Tab (`liste ouverte > ghost > mot du document > indentation`, une seule action), helpers purs partagés (`getWordFragment`, `findWordCompletions`, `normalizeGhost`, `chooseTabAction`) dans `autocomplete.js`, liste déroulante si plusieurs candidats, dropdown positionné au curseur, ghost **positionné au curseur** (fini le miroir du document, nettoyé au déplacement/scroll), complétion de mot sans espace garanti (`normalizeGhost` tronque au premier espace) et prompt `/api/ai/inline-complete` simplifié. Tests : `tests/frontend/forge-completion.test.mjs` (28). | Cause du bug : l'indentation Tab n'était pas conditionnée à l'absence de suggestion. Le ghost re-rendait tout le texte transparent + prédiction, d'où l'impression d'espaces et les erreurs d'effacement |
|
||||
| *BUG-056* | [🟡 IMPORTANT] Éditeur Forge en plein écran : l'Assistant IA s'ouvre en arrière-plan et reste invisible | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs` | Forge : passer en plein écran puis cliquer le bouton « Assistant IA » (ou `Ctrl+J`) — le panneau s'ouvre dans le document parent, masqué par l'iframe plein écran | Sortie du plein écran **avant** d'ouvrir le panneau, des deux côtés : côté iframe (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et** côté parent (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut être détenu par le document parent et non par l'iframe (dans ce cas `document.fullscreenElement` est nul dans l'iframe et sa sortie échoue). Tests : `forge-completion.test.mjs` (+1), `editor-inline.test.mjs` (+1) | Le panneau assistant est monté dans `document.body` du parent : l'API Fullscreen ne rend que l'élément plein écran et ses descendants, donc il ne peut pas s'afficher au-dessus de l'iframe Forge en plein écran. La sortie côté iframe seule ne suffisait pas quand le parent détient le plein écran |
|
||||
| *BUG-057* | [🟡 IMPORTANT] Assistant IA : le bouton « Ajouter » est inopérant dans l'éditeur Forge (fonctionne seulement dans « Editer ») | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js`, `frontend/editor-poc.html` | Ouvrir un document dans Forge, demander une réponse à l'assistant puis cliquer « Ajouter » | `_insertIntoEditor()` cible Forge (`#forge-iframe`) : `postMessage({ type: 'parent-insert', text })` ; `editor-poc.html` insère au curseur (`insertAtCursor`) et marque le tampon modifié. Repli textarea inclus. Tests : `tests/frontend/ai.test.mjs` (+3), `tests/frontend/editor-inline.test.mjs` (+1) | `state.editorView` (CodeMirror) est nul en Forge : le clic affichait « Aucun document ouvert dans l'éditeur » |
|
||||
| *BUG-058* | [🔵 MINEUR] Éditeur « Editer » : la barre de numérotation de ligne ne suit pas la couleur du thème (gutter clair `#f5f5f5` en thème sombre) | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css` | Ouvrir un document → Editer en thème sombre : la colonne des numéros de ligne reste gris clair alors que le fond de l'éditeur est sombre | Thème du gutter CodeMirror via les variables CSS (`color-mix(var(--text-primary) …)` pour le fond, `--text-secondary` pour les numéros, `--border` pour la séparation, `--text-primary` pour la ligne active) au lieu des valeurs codées en dur de CodeMirror ; test de non-régression dans `tests/frontend/editor-inline.test.mjs`. Vérifié Playwright (instance de test) : sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de` | CodeMirror applique `background:#f5f5f5` par défaut, indépendamment du thème ObsiGate ; en mode sombre le fond de l'éditeur suit `--bg-secondary` mais pas le gutter |
|
||||
| *BUG-060* | [🟡 IMPORTANT] Viewer PDF : l'affichage des pages ne fonctionne pas — seule la barre d'outils « PDF — N pages » s'affiche, le contenu reste vide | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau) | Cliquer un fichier `.pdf` dans l'arborescence | `frontend/js/viewer.js` : le rendu PDF passe de `<embed type="application/pdf">` à `<iframe>` (autorisée par `frame-src 'self'`, le stream étant same-origin). Tests : `tests/frontend/pdf-viewer.test.mjs` (+6) et `tests/e2e/pdf-viewer.spec.js` (fixture `test_vault/sample-pdf.pdf`) | Cause : la CSP durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>` → le lecteur PDF natif était bloqué (barre d'outils rendue, corps vide). Le test E2E échoue bien avec l'ancien `<embed>`. `object-src 'none'` conservé (le correctif ne désarme pas la CSP) |
|
||||
| *BUG-061* | [🟡 IMPORTANT] Assistant IA : le bouton « Plein écran » n'agrandit plus le panneau | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css`, `tests/frontend/ai.test.mjs` | Ouvrir l'assistant, redimensionner le panneau, puis cliquer « Plein écran » | La largeur du panneau est écrite en ligne par la poignée de redimensionnement / la largeur persistée (`localStorage`) ; l'inline l'emportait sur `.bookslm-panel.fullscreen { width: 100vw }`. Ajout de `!important` sur la règle plein écran. Tests : `ai.test.mjs` (+1 : classe basculée + règle CSS). Vérifié Playwright : 640 px → 1400 px (viewport) |
|
||||
| *BUG-062* | [🟡 IMPORTANT] Viewer PDF : le document ne prend pas toute la largeur quand la navigation est masquée | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css`, `tests/frontend/pdf-viewer.test.mjs`, `tests/e2e/pdf-viewer.spec.js` | Ouvrir un PDF puis masquer la barre de navigation gauche | La règle `.sidebar.hidden ~ .content-wrapper .content-area { max-width: 1200px }` (colonne de lecture centrée) s'appliquait aussi aux viewers plein cadre. Ajout de `.content-area:has(.pdf-viewer-container)` (et `.image-viewer-container`) avec `max-width: none; margin: 0`. Test E2E : `max-width` calculé = `none`, conteneur = largeur du contenu |
|
||||
| *BUG-063* | [🟡 IMPORTANT] Viewer PDF : la table des matières s'affiche mais ne navigue pas | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs`, `tests/e2e/pdf-viewer.spec.js` (fixture `test_vault/sample-pdf-toc.pdf`) | Ouvrir un PDF avec signets, puis cliquer une entrée de la TOC | Deux causes : (1) `contentWindow.location.hash='page=N'` n'atteint pas le document (lecteur PDF natif dans une fenêtre `about:blank`) ; (2) un simple changement de fragment sur `iframe.src` est une navigation same-document **ignorée** par le lecteur natif. `navigatePdfToPage()` (liens `data-page` + listeners, plus d'`onclick` inline) recharge réellement l'iframe via un paramètre de query qui change (`&_pdfpage=<ts>#page=N`). Test E2E : `src` finit par `&_pdfpage=<n>#page=3`. Vérifié en Chrome *headful* : page 1 → page 8 → page 1 (captures identiques au retour) | Le fragment seul ne suffisait pas : Chrome applique `#page=N` au **chargement**, pas lors d'un changement de fragment |
|
||||
| | | | | | | | | | | |
|
||||
|
||||
### TODOs techniques (améliorations / nouvelles tâches)
|
||||
@@ -227,6 +232,12 @@ Avant de corriger quoi que ce soit, un agent IA doit :
|
||||
| 2026-09-17 | BUG-055 (complément) | Correction | `frontend/editor-poc.html`, `frontend/js/utils.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-055 (complément)** : une complétion acceptée au `Tab` disparaissait 1–2 s plus tard. Cause : l'auto-sauvegarde (2 s) déclenche un `index_updated` SSE sur le fichier affiché, et `reloadExternalWrite` rechargeait le tampon Forge **depuis le disque**, écrasant toute frappe postérieure à la sauvegarde. Le rechargement SSE est désormais ignoré si le tampon est modifié (`parent-reload` sans `force` + `isDirty` ; garde équivalente sur le point d'auto-sauvegarde CodeMirror) ; seul `obsigate:file-written` (assistant IA) passe `force=true`. L'auto-sauvegarde ne remet plus l'état « enregistré » si des modifications sont arrivées pendant la requête (Forge + CodeMirror), et `acceptGhost()` annule la requête de prédiction en attente. Vérifié : `forge-completion.test.mjs` 31/31 (+3), `editor-inline.test.mjs` 41/41 (+1), 14 suites frontend vertes, validate-imports 38 modules. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-056 | Correction | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-056** : en plein écran Forge, l'Assistant IA s'ouvrait en arrière-plan. La sortie du plein écran est désormais faite **côté iframe** (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et côté parent** (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut appartenir au document parent (l'iframe voit alors `fullscreenElement` nul et sa sortie échoue — c'était le cas non couvert par le premier correctif). Vérifié : `forge-completion.test.mjs` 32/32 (+1), `editor-inline.test.mjs` 42/42 (+1), 14 suites frontend vertes, validate-imports 38 modules. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-057, #102 | Correction + feature | `frontend/js/bookslm.js`, `frontend/editor-poc.html`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/frontend/ai.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `docs/archive/COMPLETED_v1-v2.md`, `docs/ROADMAP.md`, `CHANGELOG.md` | **BUG-057** : le bouton « Ajouter » de l'assistant ne ciblait que `state.editorView` (CodeMirror) ; en Forge il affichait « Aucun document ouvert dans l'éditeur ». `_insertIntoEditor()` gère désormais les trois surfaces : CodeMirror, l'iframe Forge (`postMessage({ type: 'parent-insert', text })` → `insertAtCursor` dans `editor-poc.html`) et le textarea de repli. **#102** : chaque bloc de code d'une réponse reçoit un bouton « Ajouter la section » (`.bookslm-code-insert`, révélé au survol) qui insère le contenu du bloc sans les délimiteurs ` ``` `. Vérifié : `ai.test.mjs` 91/91 (+3), `editor-inline.test.mjs` 43/43 (+1), `forge-completion.test.mjs` 32/32, unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-058 | Correction | `frontend/style.css`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-058** : la barre de numérotation de ligne de l'éditeur « Editer » ne suivait pas le thème — CodeMirror peint `.cm-gutters` avec des valeurs claires codées en dur (`#f5f5f5`, bordure `#ddd`), visibles en thème sombre. Correctif : le gutter dérive des variables CSS ObsiGate (`background: color-mix(in srgb, var(--text-primary) 5%, transparent)`, `color: var(--text-secondary)`, `border-right: 1px solid var(--border)`, ligne active `color-mix(… 10% …)` / `--text-primary`), donc il suit les 15 thèmes et les 4 modes. Vérifié : `editor-inline.test.mjs` 44/44 (+1), unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0, et Playwright sur l'instance de test (route `style.css` remplacée par le fichier local) — sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de`, plus de `rgb(245,245,245)`. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-059 | Correction | `frontend/js/bookslm.js`, `tests/frontend/ai.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-059** : dans une conversation ouverte (post ancré en haut), **tout clic** dans la fenêtre de messages — lien de fichier, étapes, sélection de texte — faisait sauter toute la conversation au bas de la fenêtre. Cause : le gestionnaire `mousedown` de dépintage (prévu pour la molette/tactile/poignée de scroll) se déclenchait aussi sur un simple clic, et le retrait du padding d'ancre (`paddingBottom`) bornait le `scrollTop` à la nouvelle hauteur max → saut au bas. Correctif : helper pur `isScrollbarPress(target, clientX, container)` — un appui ne dépine que s'il vise la **poignée de scroll** (cible = conteneur + zone de gouttière droite) ; molette et tactile conservent leur comportement. Vérifié : `ai.test.mjs` 92/92 (+1), unit 9/9, validate-imports 38 modules, pytest / ruff / mypy inchangés côté backend. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-060 | Correction | `frontend/js/viewer.js`, `.gitea/workflows/ci.yml`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau), `test_vault/sample-pdf.pdf` (nouveau), `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-060** : l'ouverture d'un PDF n'affichait aucune page (barre d'outils « PDF — N pages » présente, corps vide). Cause : la CSP durcie en BUG-034 pose `object-src 'none'` — directive qui gouverne `<embed>`/`<object>` — alors que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était bloqué. Correctif : rendu dans une `<iframe>` (autorisée par `frame-src 'self'`, le stream `/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` conservé. Tests : `pdf-viewer.test.mjs` 6/6 (statique : pas d'`<embed>`, CSP `frame-src 'self'`, iframe pleine hauteur), `pdf-viewer.spec.js` (E2E : iframe + stream `application/pdf` 200/206 + zéro violation CSP ; échoue bien avec l'ancien `<embed>`). Vérifié : pytest 1184 passed / 6 skipped, frontend 14 suites JSDOM vertes, validate-imports 38 modules, ruff/mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-035, BUG-036, BUG-037, BUG-038, BUG-039, BUG-040 | Correction | `backend/secret_redactor.py`, `backend/collab.py`, `backend/auth/{middleware,password,router}.py`, `backend/indexer.py`, `backend/main.py`, `tests/test_api_main.py`, `tests/test_auth.py`, `tests/test_auth_api.py`, `tests/test_collab.py`, `tests/test_pdf.py`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **Lot de 6 bugs mineurs (P2/P3)** : BUG-035 masquage hex conditionné au contexte (git/SHA épargnés) ; BUG-036 jeton WebSocket cookie-only (plus de `?token=`) + plafond de trame 16 Mio ; BUG-037 garde-fou au démarrage (refus d'un bind public sans auth sauf `OBSIGATE_ALLOW_INSECURE=true`) ; BUG-038 Argon2 recalibré 19 Mio/t=2/p=1 ; BUG-039 login uniforme 401 (fini 429/403 distinctifs) ; BUG-040 extraction PDF différée via `enrich_pdf_texts()`. Vérifié : pytest 1204 passed / 6 skipped, ruff 0, mypy 0 (77 fichiers), frontend validate-imports 38 modules + unit 9/9. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-061, BUG-062, BUG-063 | Correction | `frontend/style.css`, `frontend/js/viewer.js`, `tests/frontend/ai.test.mjs`, `tests/frontend/pdf-viewer.test.mjs`, `tests/e2e/pdf-viewer.spec.js`, `test_vault/sample-pdf-toc.pdf` (nouveau), `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **Viewer PDF & assistant IA** : BUG-061 le bouton plein écran du panneau assistant l'emportait mal sur la largeur inline (redimensionnement/persistée) → `width: 100vw !important` ; BUG-062 le plafond de lecture 1200 px s'appliquait au PDF quand la navigation était masquée → `:has(.pdf-viewer-container)` en `max-width:none` ; BUG-063 la TOC PDF ne naviguait pas (`contentWindow` = `about:blank`) → `navigatePdfToPage()` recharge l'iframe avec `#page=N`, liens `data-page` sans `onclick` inline. Vérifié : `ai.test.mjs` 93/93, `pdf-viewer.test.mjs` 8/8, validate-imports 38 modules (311 exports), unit 9/9, E2E `pdf-viewer.spec.js` 3/3, et Playwright sur l'instance de test (plein écran 640→1400 px, `src` → `#page=3`, `max-width:none`). | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-17 | BUG-063 (complément) | Correction | `frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs`, `tests/e2e/pdf-viewer.spec.js`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-063 non résolu au premier correctif** : définir `iframe.src = base + '#page=N'` ne change que le fragment → navigation same-document que le lecteur PDF natif ignore. Diagnostic en Chrome *headful* (comparaison de captures) : fragment présent au chargement = OK ; changement de fragment après chargement = aucun effet ; changement de query + fragment = OK. `navigatePdfToPage()` ajoute donc un paramètre de query horodaté (`&_pdfpage=<ts>#page=N`) pour forcer un vrai rechargement. Vérifié via l'UI de l'app (Chrome headful) : page 1 → page 8 → retour page 1 (hash de capture identique au retour). Tests : `pdf-viewer.test.mjs` 8/8, E2E `pdf-viewer.spec.js` 3/3. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
|
||||
---
|
||||
|
||||
|
||||
+6
-22
@@ -1,6 +1,6 @@
|
||||
# ObsiGate — Roadmap
|
||||
|
||||
> **Version :** 2.9.0 | **Dernière mise à jour :** 2026-09-17
|
||||
> **Version :** 2.11.5 | **Dernière mise à jour :** 2026-09-17
|
||||
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
|
||||
> vers les fonctionnalités livrées.
|
||||
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
|
||||
@@ -78,23 +78,6 @@
|
||||
- [x] Personnalisation (clé à molette) : ajouter / supprimer / réordonner les commandes
|
||||
- [x] i18n FR/EN + tests frontend (helpers purs) + E2E mobile
|
||||
|
||||
### 92. Assistant IA — Écosystème d'outils (phase 2 : web étendu, sources connectées, documents)
|
||||
|
||||
- **Effort :** 3-5 jours | **Impact :** 🟠 | **Zone :** backend (`backend/tools/`)
|
||||
- **Dépend de :** #91 (registre + section « steps » + `web_search`/`fetch_url` livrés)
|
||||
- **Description :** étendre le catalogue d'outils de l'assistant au-delà du vault, en
|
||||
suivant la feuille de route technique détaillée :
|
||||
[features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) (frameworks évalués,
|
||||
bibliothèques par catégorie, transverse retry/cache/secrets/async).
|
||||
- **Sous-tâches :**
|
||||
- [x] `web_search` : chaîne de repli sans clé (SearXNG → DuckDuckGo → Bing, `OBSIGATE_WEB_FALLBACK`) — BUG-051
|
||||
- [ ] `web_search` : fournisseurs optionnels à clé (Tavily, Brave, SerpAPI, Exa)
|
||||
- [ ] `fetch_url` : pages dynamiques via Playwright (worker isolé) ; crawl multi-pages Scrapy en tâche de fond
|
||||
- [ ] Sources connectées : Gitea/GitHub (priorité haute) puis Google Drive / OneDrive (OAuth2 `authlib`)
|
||||
- [ ] Production de documents : conversion, tableurs, PDF/Word (outils WRITE + confirmation)
|
||||
- [ ] Transverse : `tenacity` (backoff), cache SQLite des résultats web avec TTL, secrets via Infisical
|
||||
- [ ] Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests (httpx mocké)
|
||||
|
||||
---
|
||||
|
||||
## ⚪ Backlog — Sécurité, architecture & performance (P0/P1)
|
||||
@@ -201,6 +184,8 @@
|
||||
| 101 | Forge — Assistant IA partagé (bouton AI Panel = assistant, fournisseur/modèle configuré, autocomplétion) + plein écran Forge/Editer | 2.8.0 | [features/forge-assistant.md](./features/forge-assistant.md) |
|
||||
| BUG-057 | Assistant IA — bouton « Ajouter » fonctionnel dans l'éditeur Forge (en plus d'« Editer ») | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
|
||||
| 102 | Assistant IA — bouton « Ajouter la section » par bloc de code (insertion du bloc seul) | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
|
||||
| 92 | Assistant IA — Écosystème d'outils phase 2 (recherche à clé, cache/retry, Playwright, crawl, Gitea/GitHub, documents XLSX/DOCX/CSV/PDF) | 2.10.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
|
||||
| 103 | Configuration — clés utilisateur des sources connectées & recherche à clé (page Configurations, `data/api_keys.json`, priorité sur l'env) | 2.11.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
|
||||
|
||||
---
|
||||
|
||||
@@ -208,12 +193,11 @@
|
||||
|
||||
| Priorité | Items | Effort total estimé |
|
||||
|---|---|---|
|
||||
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100 | ~114 jours réalisés |
|
||||
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100, #102, #92 | ~114 jours réalisés |
|
||||
| 🔵 P2 restant | #77 Desktop : signature de code (non retenue), 6 tests E2E **manuels** ([protocole](./DESKTOP_E2E_CHECKLIST.md)) | ~0,5-1 jour |
|
||||
| ⚪ P4 restant | #73 Sync (6-8j) | 6-8 jours |
|
||||
| ⚪ P2 restant | #92 Assistant IA — écosystème d'outils phase 2 (web étendu, sources connectées, documents) | 3-5 jours |
|
||||
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (issues BUG-035 → BUG-040) | ~15-23 jours |
|
||||
| **Total restant** | **10 items + finitions** | **~30-47 jours** |
|
||||
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (BUG-035 → BUG-040 corrigés) | ~15-23 jours |
|
||||
| **Total restant** | **7 items + finitions** | **~27-42 jours** |
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# #92 — Assistant IA — Écosystème d'outils : feuille de route technique
|
||||
|
||||
> **Statut :** ⚪ Backlog (phase 1 livrée dans #91)
|
||||
> **Statut :** ✅ livré (phase 2, version 2.10.0) — phase 1 livrée dans #91
|
||||
> **Effort estimé :** 3-5 jours pour la phase 2 | **Impact :** 🟠
|
||||
> **Références :** [Roadmap](../ROADMAP.md) · [Outils & MCP #79](./ai-tools-mcp.md) ·
|
||||
> [Fenêtre de discussion #91](./ai-assistant-conversation-ux.md) · [Changelog](../../CHANGELOG.md)
|
||||
@@ -36,6 +36,17 @@ réécriture de la boucle n'est nécessaire.
|
||||
|
||||
## 3. Phase 2 — catégories à implémenter
|
||||
|
||||
> **Livré (2.10.0, #92).** Récapitulatif des décisions finales :
|
||||
|
||||
| Catégorie | Décision livrée |
|
||||
|---|---|
|
||||
| Recherche web étendue | Tavily, Brave, SerpAPI, Exa à clé (`OBSIGATE_*_API_KEY`), essayés avant SearXNG ; ordre via `OBSIGATE_WEB_PROVIDERS` |
|
||||
| Lecture de pages | `fetch_url(render=True)` → worker Playwright isolé (`backend/tools/webrender.py`), dépendance optionnelle + erreur explicite |
|
||||
| Crawl multi-pages | `crawl_site` (WRITE + confirmation) : BFS httpx borné (≤ 20 pages, même hôte, SSRF sur chaque URL) → condensé Markdown dans le vault. Scrapy écarté (dépendance lourde inutile à cette échelle) |
|
||||
| Sources connectées | Gitea + GitHub (`git_list_repos`, `git_search_issues`, `git_get_file`) via env/Infisical ; drives cloud (Drive/OneDrive) orientés serveur MCP externe (#79). **#103 (2.11.0)** : les clés (URL Gitea, tokens Gitea/GitHub, clés Tavily/Brave/SerpAPI/Exa) se saisissent aussi dans la page Configurations — `backend/tools/secrets.py`, valeur stockée prioritaire sur l'env |
|
||||
| Production de documents | `create_xlsx`, `create_docx`, `create_csv`, `create_pdf` — WRITE + confirmation, écrit via `save_raw_file(allow_docs=True)` (path safety + backup) |
|
||||
| Transverse | Cache SQLite (`webcache.py`, TTL `OBSIGATE_WEB_CACHE_TTL`), retry backoff maison (`OBSIGATE_WEB_RETRY`), secrets par env (Infisical-compatible) |
|
||||
|
||||
### 3.1 Recherche web étendue (`web_search`)
|
||||
- **Fallback sans clé — ✅ livré (BUG-051)** : chaîne de fournisseurs dans
|
||||
`backend/tools/web.py` — SearXNG auto-hébergé (`OBSIGATE_SEARXNG_URL`) puis, si
|
||||
|
||||
@@ -109,11 +109,14 @@ Serveur → client :
|
||||
|
||||
## Sécurité
|
||||
|
||||
- Authentification obligatoire si `OBSIGATE_AUTH_ENABLED=true` (cookie ou `?token=`).
|
||||
- Authentification obligatoire si `OBSIGATE_AUTH_ENABLED=true` : le jeton est lu depuis le cookie
|
||||
HttpOnly `access_token` (envoyé lors du handshake same-origin). Le jeton en query string
|
||||
(`?token=`) n'est **plus accepté** (BUG-036 : URLs journalisées par les proxies).
|
||||
- Vérification `check_vault_access()` par connexion (un utilisateur ne peut pas rejoindre une room
|
||||
d'une vault non autorisée).
|
||||
- `resolve_safe_path()` empêche toute traversée de chemin (`../../`).
|
||||
- Bornes anti-abus : `MAX_UPDATE_BYTES` (8 Mo) par mise à jour, `MAX_TEXT_CHARS` (8 Mio) par snapshot.
|
||||
- Bornes anti-abus : `MAX_UPDATE_BYTES` (8 Mo) par mise à jour, `MAX_TEXT_CHARS` (8 Mio) par snapshot,
|
||||
`MAX_MESSAGE_CHARS` (16 Mio) par trame brute.
|
||||
- Le serveur ne décode pas le binaire Yjs : il le stocke et le relaie tel quel (pas de surface
|
||||
d'attaque supplémentaire côté parsing).
|
||||
|
||||
|
||||
@@ -8,6 +8,7 @@
|
||||
|
||||
- **Implémentation réelle (vérifiée 2026-09-07) :**
|
||||
- **Bugs corrigés (2026-09) :** `api_pdf_stream` crashait en 500 (`NameError: current_user` jamais injecté) ; l'indexation incrémentale du watcher faisait `read_text()` sur les PDFs (garbage) ; Range/206 et `pdf/info` absents malgré le texte ci-dessous.
|
||||
- **BUG-060 (2026-09-17) :** l'affichage inline ne fonctionnait plus — la CSP durcie en BUG-034 (`object-src 'none'`) bloquait l'`<embed>` du viewer (barre d'outils rendue, corps vide). Le rendu passe par une `<iframe>` (autorisée par `frame-src 'self'`), conforme à E1. Tests : `tests/frontend/pdf-viewer.test.mjs` + `tests/e2e/pdf-viewer.spec.js`.
|
||||
- `GET /api/file/{vault}/pdf/info` — métadonnées seules sans transférer le document (C3)
|
||||
- Stream avec `Accept-Ranges` + 206 Partial Content (single range, suffix-range, 416) (C2)
|
||||
- `OBSIGATE_PDF_MAX_SIZE_MB` (50) + `OBSIGATE_PDF_EXTRACT_TIMEOUT` (30s via thread-pool) (B4/G3)
|
||||
|
||||
@@ -1543,6 +1543,7 @@
|
||||
<li><a href="#cfg-hidden-files" class="help-nav-link" data-i18n="config.section_hidden"></a></li>
|
||||
<li><a href="#cfg-diags" class="help-nav-link" data-i18n="settings.diagnostics"></a></li>
|
||||
<li><a href="#cfg-ai" class="help-nav-link" data-i18n="settings.ai"></a></li>
|
||||
<li><a href="#cfg-sources" class="help-nav-link" data-i18n="config.section_sources">Sources connectées</a></li>
|
||||
<li><a href="#cfg-themes" class="help-nav-link" data-i18n="settings.themes"></a></li>
|
||||
<li><a href="#cfg-profile" class="help-nav-link" data-i18n="settings.profile"></a></li>
|
||||
<li><a href="#cfg-security" class="help-nav-link" data-i18n="settings.security"></a></li>
|
||||
@@ -2260,6 +2261,133 @@
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- Connected sources & keyed web search (#103) -->
|
||||
<section
|
||||
class="config-section help-section"
|
||||
id="cfg-sources"
|
||||
>
|
||||
<h2 data-i18n="config.section_sources"
|
||||
>Sources connectées & recherche web</h2
|
||||
>
|
||||
<p
|
||||
class="config-description"
|
||||
data-i18n="config.sources_desc"
|
||||
>
|
||||
Clés utilisées par les outils de l'Assistant IA
|
||||
(recherche web à clé, Gitea, GitHub). Elles sont
|
||||
stockées localement et priment sur les variables
|
||||
d'environnement.
|
||||
</p>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-tavily-key"
|
||||
>Tavily API Key</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-tavily-key"
|
||||
class="config-input"
|
||||
placeholder="tvly-..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-brave-key"
|
||||
>Brave Search API Key</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-brave-key"
|
||||
class="config-input"
|
||||
placeholder="BSA..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-serpapi-key"
|
||||
>SerpAPI Key</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-serpapi-key"
|
||||
class="config-input"
|
||||
placeholder="..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-exa-key"
|
||||
>Exa API Key</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-exa-key"
|
||||
class="config-input"
|
||||
placeholder="..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-gitea-url"
|
||||
data-i18n="config.gitea_url"
|
||||
>URL Gitea</label
|
||||
>
|
||||
<input
|
||||
type="text"
|
||||
id="cfg-gitea-url"
|
||||
class="config-input"
|
||||
placeholder="https://git.example.net"
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-gitea-token"
|
||||
>Gitea Token</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-gitea-token"
|
||||
class="config-input"
|
||||
placeholder="token personnel..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div class="config-row">
|
||||
<label
|
||||
class="config-label"
|
||||
for="cfg-github-token"
|
||||
>GitHub Token</label
|
||||
>
|
||||
<input
|
||||
type="password"
|
||||
id="cfg-github-token"
|
||||
class="config-input"
|
||||
placeholder="ghp_..."
|
||||
autocomplete="off"
|
||||
/>
|
||||
</div>
|
||||
<div
|
||||
class="config-actions-row"
|
||||
style="margin-top: 16px"
|
||||
>
|
||||
<button class="config-btn-save"
|
||||
id="cfg-save-tool-keys" data-i18n="help.shortcut_save">
|
||||
Sauvegarder
|
||||
</button>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- Themes -->
|
||||
<section
|
||||
class="config-section help-section"
|
||||
|
||||
+25
-1
@@ -43,6 +43,25 @@ const PANEL_MIN_WIDTH = 320;
|
||||
const PANEL_MAX_WIDTH = 1000;
|
||||
const PANEL_WIDTH_KEY = 'obsigate-bookslm-width';
|
||||
|
||||
/**
|
||||
* BUG-059 — Is this pointer press a scrollbar drag (and only that)?
|
||||
*
|
||||
* Only genuine scroll gestures may release the top-pinning of the latest
|
||||
* question (wheel / touchmove / scrollbar drag). A scrollbar drag reports the
|
||||
* scrollable container itself as the event target and lands inside the
|
||||
* vertical scrollbar gutter. A plain click anywhere in the content — a link,
|
||||
* a file path, the steps toggle, text selection — must NOT unpin: clearing
|
||||
* the anchor padding would clamp the scroll position and jump the whole
|
||||
* thread to the bottom of the conversation window.
|
||||
*/
|
||||
export function isScrollbarPress(target, clientX, container) {
|
||||
if (!container || target !== container) return false;
|
||||
const rect = container.getBoundingClientRect();
|
||||
if (!rect || !Number.isFinite(rect.right)) return false;
|
||||
const GUTTER = 24; // conservative vertical-scrollbar width estimate
|
||||
return clientX >= rect.right - GUTTER;
|
||||
}
|
||||
|
||||
/**
|
||||
* Accent- and case-insensitive normalization used to match `/` commands and
|
||||
* `@` mentions: skill ids/labels and vault paths may contain accented
|
||||
@@ -985,7 +1004,12 @@ class BooksLM {
|
||||
};
|
||||
messagesEl.addEventListener('wheel', unpin, { passive: true });
|
||||
messagesEl.addEventListener('touchmove', unpin, { passive: true });
|
||||
messagesEl.addEventListener('mousedown', unpin, { passive: true });
|
||||
// BUG-059: a scrollbar drag only — a plain click on the content must
|
||||
// not unpin (it would clear the anchor padding and jump the thread to
|
||||
// the bottom of the window).
|
||||
messagesEl.addEventListener('mousedown', (e) => {
|
||||
if (isScrollbarPress(e.target, e.clientX, messagesEl)) unpin();
|
||||
}, { passive: true });
|
||||
}
|
||||
|
||||
// Close the session / command menus when clicking elsewhere in the panel.
|
||||
|
||||
@@ -715,6 +715,7 @@ function initConfigModal() {
|
||||
await loadHiddenFilesSettings();
|
||||
loadWebhooksUI();
|
||||
loadSharesUI();
|
||||
loadToolKeys();
|
||||
safeCreateIcons();
|
||||
});
|
||||
|
||||
@@ -765,6 +766,9 @@ function initConfigModal() {
|
||||
if (saveAIKeysBtn) saveAIKeysBtn.addEventListener("click", saveAIKeys);
|
||||
const testAIKeysBtn = document.getElementById("cfg-test-ai-keys");
|
||||
if (testAIKeysBtn) testAIKeysBtn.addEventListener("click", testAIKeys);
|
||||
// Tool & connected-source keys (#103)
|
||||
const saveToolKeysBtn = document.getElementById("cfg-save-tool-keys");
|
||||
if (saveToolKeysBtn) saveToolKeysBtn.addEventListener("click", saveToolKeys);
|
||||
// Default provider/model selection
|
||||
const aiDefaultProviderSel = document.getElementById("cfg-ai-default-provider");
|
||||
if (aiDefaultProviderSel) {
|
||||
@@ -1715,6 +1719,102 @@ async function testAIKeys() {
|
||||
}
|
||||
|
||||
|
||||
// ── Tool & connected-source keys (#103) ──
|
||||
const TOOL_KEY_MAP = {
|
||||
"cfg-tavily-key": "OBSIGATE_TAVILY_API_KEY",
|
||||
"cfg-brave-key": "OBSIGATE_BRAVE_API_KEY",
|
||||
"cfg-serpapi-key": "OBSIGATE_SERPAPI_API_KEY",
|
||||
"cfg-exa-key": "OBSIGATE_EXA_API_KEY",
|
||||
"cfg-gitea-url": "OBSIGATE_GITEA_URL",
|
||||
"cfg-gitea-token": "OBSIGATE_GITEA_TOKEN",
|
||||
"cfg-github-token": "OBSIGATE_GITHUB_TOKEN",
|
||||
};
|
||||
|
||||
function _ensureToolKeyUI() {
|
||||
for (const [inputId] of Object.entries(TOOL_KEY_MAP)) {
|
||||
const input = document.getElementById(inputId);
|
||||
if (!input) continue;
|
||||
const row = input.closest(".config-row");
|
||||
if (!row || row.dataset.toolKeyEnhanced) continue;
|
||||
row.dataset.toolKeyEnhanced = "1";
|
||||
row.style.cssText += "display:flex;align-items:center;gap:8px;flex-wrap:wrap;";
|
||||
const badge = document.createElement("span");
|
||||
badge.id = inputId + "-badge";
|
||||
badge.style.cssText = "font-size:11px;padding:2px 8px;border-radius:10px;white-space:nowrap;";
|
||||
row.appendChild(badge);
|
||||
const delBtn = document.createElement("button");
|
||||
delBtn.type = "button";
|
||||
delBtn.id = inputId + "-delete";
|
||||
delBtn.className = "config-btn-secondary";
|
||||
delBtn.style.cssText = "font-size:11px;padding:4px 10px;color:var(--danger,#e74c3c);border-color:var(--danger,#e74c3c);cursor:pointer;display:none;";
|
||||
delBtn.textContent = "\u00d7 " + t("config.delete_key");
|
||||
delBtn.addEventListener("click", () => deleteToolKey(inputId));
|
||||
row.appendChild(delBtn);
|
||||
}
|
||||
}
|
||||
|
||||
function _setToolKeyBadge(inputId, hasKey) {
|
||||
const badge = document.getElementById(inputId + "-badge");
|
||||
const delBtn = document.getElementById(inputId + "-delete");
|
||||
if (badge) {
|
||||
if (hasKey) {
|
||||
badge.textContent = "\u2713 " + t("config.key_set");
|
||||
badge.style.background = "var(--success-bg, #27ae6022)";
|
||||
badge.style.color = "var(--success, #27ae60)";
|
||||
badge.style.border = "1px solid var(--success, #27ae60)";
|
||||
} else {
|
||||
badge.textContent = t("config.key_unset");
|
||||
badge.style.background = "var(--muted-bg, #ffffff10)";
|
||||
badge.style.color = "var(--text-muted, #888)";
|
||||
badge.style.border = "1px solid var(--border, #444)";
|
||||
}
|
||||
}
|
||||
if (delBtn) delBtn.style.display = hasKey ? "inline-block" : "none";
|
||||
}
|
||||
|
||||
async function loadToolKeys() {
|
||||
_ensureToolKeyUI();
|
||||
try {
|
||||
const data = await api("/api/config/tool-keys");
|
||||
for (const [inputId, name] of Object.entries(TOOL_KEY_MAP)) {
|
||||
const input = document.getElementById(inputId);
|
||||
const val = data[name] || "";
|
||||
if (input && !input.value.trim()) input.placeholder = val || input.placeholder;
|
||||
_setToolKeyBadge(inputId, !!val);
|
||||
}
|
||||
} catch(e) { /* non-admin: the section stays inert */ }
|
||||
}
|
||||
|
||||
async function saveToolKeys() {
|
||||
const keys = {};
|
||||
for (const [id, name] of Object.entries(TOOL_KEY_MAP)) {
|
||||
const input = document.getElementById(id);
|
||||
if (input && input.value.trim()) keys[name] = input.value.trim();
|
||||
}
|
||||
if (!Object.keys(keys).length) {
|
||||
showToast(t("config.no_keys"), "warning");
|
||||
return;
|
||||
}
|
||||
try {
|
||||
await api("/api/config/tool-keys", { method: "POST", body: JSON.stringify(keys) });
|
||||
showToast(t("config.api_keys_saved"), "success");
|
||||
Object.keys(TOOL_KEY_MAP).forEach(id => { const el = document.getElementById(id); if (el) el.value = ""; });
|
||||
loadToolKeys();
|
||||
} catch(e) { showToast("Erreur: " + e.message, "error"); }
|
||||
}
|
||||
|
||||
async function deleteToolKey(inputId) {
|
||||
const name = TOOL_KEY_MAP[inputId];
|
||||
if (!name) return;
|
||||
if (!confirm(t("config.delete_key_confirm") + " " + name + " ?")) return;
|
||||
try {
|
||||
await api("/api/config/tool-keys/" + name, { method: "DELETE" });
|
||||
showToast(t("config.key_deleted") + " " + name, "success");
|
||||
loadToolKeys();
|
||||
} catch(e) { showToast("Erreur: " + e.message, "error"); }
|
||||
}
|
||||
|
||||
|
||||
export {
|
||||
initSidebarTabs,
|
||||
initConfigModal,
|
||||
|
||||
+30
-2
@@ -518,6 +518,28 @@ function applyPrettyHighlight(codeEl, lang, text) {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Jump the inline PDF viewer to a specific page.
|
||||
*
|
||||
* The browser's built-in PDF viewer lives in an ``about:blank`` content window
|
||||
* (so ``contentWindow.location.hash`` never reaches the document) **and** it
|
||||
* ignores a same-document fragment navigation: changing only ``#page=N`` on the
|
||||
* iframe ``src`` does not move the page. A changing query parameter forces a
|
||||
* real reload, and the ``#page=N`` fragment is then honoured at load — the only
|
||||
* reliable way to target a page with the native viewer.
|
||||
*
|
||||
* @param {HTMLElement} area - Content area containing the ``.pdf-iframe``.
|
||||
* @param {string|number} page - 1-based page number from the PDF outline.
|
||||
*/
|
||||
export function navigatePdfToPage(area, page) {
|
||||
const iframe = area && area.querySelector('.pdf-iframe');
|
||||
if (!iframe || page === null || page === undefined || page === '') return;
|
||||
const base = iframe.getAttribute('data-pdf-url') || iframe.src.split('#')[0];
|
||||
iframe.setAttribute('data-pdf-url', base);
|
||||
const sep = base.includes('?') ? '&' : '?';
|
||||
iframe.src = `${base}${sep}_pdfpage=${Date.now()}#page=${page}`;
|
||||
}
|
||||
|
||||
export function renderFile(data) {
|
||||
// #93 — An inline edition session (#editor-container mounted in the content
|
||||
// area) is destroyed by this very re-render: release it first so the editor
|
||||
@@ -537,7 +559,7 @@ export function renderFile(data) {
|
||||
tocHtml = '<div class="pdf-toc"><h3>Table des matières</h3><ul>';
|
||||
for (const item of toc) {
|
||||
const indent = (item.level - 1) * 16;
|
||||
tocHtml += `<li style="padding-left:${indent}px"><a href="#" onclick="document.querySelector('.pdf-iframe').contentWindow.location.hash='page=${item.page}';return false">${escapeHtml(item.title)}</a> <span class="toc-page">p.${item.page}</span></li>`;
|
||||
tocHtml += `<li style="padding-left:${indent}px"><a href="#" data-page="${item.page}">${escapeHtml(item.title)}</a> <span class="toc-page">p.${item.page}</span></li>`;
|
||||
}
|
||||
tocHtml += '</ul></div>';
|
||||
}
|
||||
@@ -555,9 +577,15 @@ export function renderFile(data) {
|
||||
</div>
|
||||
<div class="pdf-body">
|
||||
${tocHtml}
|
||||
<embed src="${pdfUrl}" class="pdf-iframe" type="application/pdf" title="${escapeHtml(data.title)}"></embed>
|
||||
<iframe src="${pdfUrl}" data-pdf-url="${pdfUrl}" class="pdf-iframe" title="${escapeHtml(data.title)}"></iframe>
|
||||
</div>
|
||||
</div>`;
|
||||
area.querySelectorAll('.pdf-toc a[data-page]').forEach((link) => {
|
||||
link.addEventListener('click', (e) => {
|
||||
e.preventDefault();
|
||||
navigatePdfToPage(area, link.getAttribute('data-page'));
|
||||
});
|
||||
});
|
||||
lucide.createIcons();
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -364,6 +364,14 @@
|
||||
"config.ai_model": "Model",
|
||||
"config.ai_openrouter_label": "OpenRouter API Key",
|
||||
"config.api_keys_saved": "API keys saved",
|
||||
"config.section_sources": "🔗 Connected sources & search",
|
||||
"config.sources_desc": "Keys used by the AI assistant tools (keyed web search: Tavily, Brave, SerpAPI, Exa; connected sources: Gitea, GitHub). They are stored server-side and take precedence over environment variables.",
|
||||
"config.gitea_url": "Gitea URL",
|
||||
"config.key_set": "Configured",
|
||||
"config.key_unset": "Not configured",
|
||||
"config.delete_key": "Delete",
|
||||
"config.delete_key_confirm": "Delete key",
|
||||
"config.key_deleted": "Key deleted:",
|
||||
"config.backups": "Backups",
|
||||
"config.backups_desc": "Manage automatic file backups.",
|
||||
"config.client_config": "Client config",
|
||||
@@ -1822,6 +1830,14 @@
|
||||
"ai.step.vaults": "Listed the vaults",
|
||||
"ai.step.fetch_url": "Opened a web page: {value}",
|
||||
"ai.step.web_search": "Searched the web: {value}",
|
||||
"ai.step.crawl": "Crawled a site: {value}",
|
||||
"ai.step.git_repos": "Listed repositories ({value})",
|
||||
"ai.step.git_issues": "Searched issues: {value}",
|
||||
"ai.step.git_file": "Read a repo file: {value}",
|
||||
"ai.step.xlsx_create": "Spreadsheet proposed: {value}",
|
||||
"ai.step.docx_create": "Word document proposed: {value}",
|
||||
"ai.step.csv_create": "CSV file proposed: {value}",
|
||||
"ai.step.pdf_create": "PDF document proposed: {value}",
|
||||
"bookslm.copied": "Copied to clipboard",
|
||||
"bookslm.error": "AI service error",
|
||||
"bookslm.regenerate": "Regenerate",
|
||||
|
||||
@@ -364,6 +364,14 @@
|
||||
"config.ai_model": "Modèle",
|
||||
"config.ai_openrouter_label": "OpenRouter API Key",
|
||||
"config.api_keys_saved": "Clés API sauvegardées",
|
||||
"config.section_sources": "🔗 Sources connectées & recherche",
|
||||
"config.sources_desc": "Clés utilisées par les outils de l'Assistant IA (recherche web à clé : Tavily, Brave, SerpAPI, Exa ; sources connectées : Gitea, GitHub). Elles sont stockées sur le serveur et priment sur les variables d'environnement.",
|
||||
"config.gitea_url": "URL Gitea",
|
||||
"config.key_set": "Configuré",
|
||||
"config.key_unset": "Non configuré",
|
||||
"config.delete_key": "Supprimer",
|
||||
"config.delete_key_confirm": "Supprimer la clé",
|
||||
"config.key_deleted": "Clé supprimée :",
|
||||
"config.backups": "Sauvegardes",
|
||||
"config.backups_desc": "Gérez les sauvegardes automatiques de vos fichiers.",
|
||||
"config.client_config": "Configuration client",
|
||||
@@ -1822,6 +1830,14 @@
|
||||
"ai.step.vaults": "Liste des vaults consultée",
|
||||
"ai.step.fetch_url": "Page web consultée : {value}",
|
||||
"ai.step.web_search": "Recherche sur le web : {value}",
|
||||
"ai.step.crawl": "Site exploré : {value}",
|
||||
"ai.step.git_repos": "Dépôts listés ({value})",
|
||||
"ai.step.git_issues": "Issues recherchées : {value}",
|
||||
"ai.step.git_file": "Fichier de dépôt lu : {value}",
|
||||
"ai.step.xlsx_create": "Tableur proposé : {value}",
|
||||
"ai.step.docx_create": "Document Word proposé : {value}",
|
||||
"ai.step.csv_create": "Fichier CSV proposé : {value}",
|
||||
"ai.step.pdf_create": "Document PDF proposé : {value}",
|
||||
"bookslm.copied": "Réponse copiée dans le presse-papiers",
|
||||
"bookslm.error": "Erreur du service AI",
|
||||
"bookslm.regenerate": "Régénérer",
|
||||
|
||||
+25
-1
@@ -1473,6 +1473,14 @@ select {
|
||||
margin: 0 auto;
|
||||
max-width: 1200px;
|
||||
}
|
||||
|
||||
/* Full-bleed viewers (PDF, images) must use the whole width when the
|
||||
navigation sidebar is hidden instead of the centered reading column. */
|
||||
.sidebar.hidden ~ .content-wrapper .content-area:has(.pdf-viewer-container),
|
||||
.sidebar.hidden ~ .content-wrapper .content-area:has(.image-viewer-container) {
|
||||
margin: 0;
|
||||
max-width: none;
|
||||
}
|
||||
.content-area::-webkit-scrollbar {
|
||||
width: 8px;
|
||||
}
|
||||
@@ -3592,6 +3600,20 @@ select {
|
||||
overflow-x: auto !important;
|
||||
max-width: 100%;
|
||||
}
|
||||
/* BUG-058: CodeMirror paints its line-number gutter with hardcoded light
|
||||
defaults (#f5f5f5 background, #ddd border), so the bar stayed pale in dark
|
||||
themes while the editor body followed the theme. Deriving the tint from
|
||||
`--text-primary` keeps the gutter subtle and makes it adapt to every theme
|
||||
and mode (dark, light, high-contrast, sepia). */
|
||||
.cm-editor .cm-gutters {
|
||||
background: color-mix(in srgb, var(--text-primary) 5%, transparent) !important;
|
||||
color: var(--text-secondary) !important;
|
||||
border-right: 1px solid var(--border) !important;
|
||||
}
|
||||
.cm-editor .cm-gutters .cm-activeLineGutter {
|
||||
background: color-mix(in srgb, var(--text-primary) 10%, transparent) !important;
|
||||
color: var(--text-primary) !important;
|
||||
}
|
||||
.fallback-editor {
|
||||
width: 100%;
|
||||
min-height: 100%;
|
||||
@@ -9478,7 +9500,9 @@ body.popup-mode .content-area {
|
||||
.bookslm-toolbar .ai-picker { margin-left: 0; padding-left: 0; border-left: none;
|
||||
flex-wrap: wrap; row-gap: 4px; }
|
||||
.bookslm-toolbar .ai-picker select { max-width: 220px; }
|
||||
.bookslm-panel.fullscreen { width: 100vw; }
|
||||
/* `!important` is required: the panel width is also written inline by the
|
||||
resize handle / persisted width, and an inline style would otherwise win. */
|
||||
.bookslm-panel.fullscreen { width: 100vw !important; }
|
||||
.bookslm-status { padding: 6px 16px; font-size: 12px; color: var(--text-secondary);
|
||||
border-bottom: 1px solid var(--border); display: flex; gap: 12px; align-items: center; }
|
||||
.bookslm-status .bookslm-status-text { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "obsigate",
|
||||
"version": "2.9.0",
|
||||
"version": "2.11.5",
|
||||
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
|
||||
"main": "patch.js",
|
||||
"directories": {
|
||||
|
||||
@@ -0,0 +1,130 @@
|
||||
%PDF-1.3
|
||||
%“Œ‹ž ReportLab Generated PDF document (opensource)
|
||||
1 0 obj
|
||||
<<
|
||||
/F1 2 0 R
|
||||
>>
|
||||
endobj
|
||||
2 0 obj
|
||||
<<
|
||||
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
|
||||
>>
|
||||
endobj
|
||||
3 0 obj
|
||||
<<
|
||||
/Contents 13 0 R /MediaBox [ 0 0 612 792 ] /Parent 12 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
4 0 obj
|
||||
<<
|
||||
/Contents 14 0 R /MediaBox [ 0 0 612 792 ] /Parent 12 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
5 0 obj
|
||||
<<
|
||||
/Contents 15 0 R /MediaBox [ 0 0 612 792 ] /Parent 12 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
6 0 obj
|
||||
<<
|
||||
/Outlines 8 0 R /PageMode /UseNone /Pages 12 0 R /Type /Catalog
|
||||
>>
|
||||
endobj
|
||||
7 0 obj
|
||||
<<
|
||||
/Author (anonymous) /CreationDate (D:20260917204040-04'00') /Creator (anonymous) /Keywords () /ModDate (D:20260917204040-04'00') /Producer (ReportLab PDF Library - \(opensource\))
|
||||
/Subject (unspecified) /Title (untitled) /Trapped /False
|
||||
>>
|
||||
endobj
|
||||
8 0 obj
|
||||
<<
|
||||
/Count 3 /First 9 0 R /Last 11 0 R /Type /Outlines
|
||||
>>
|
||||
endobj
|
||||
9 0 obj
|
||||
<<
|
||||
/Dest [ 3 0 R /Fit ] /Next 10 0 R /Parent 8 0 R /Title (Page One)
|
||||
>>
|
||||
endobj
|
||||
10 0 obj
|
||||
<<
|
||||
/Dest [ 4 0 R /Fit ] /Next 11 0 R /Parent 8 0 R /Prev 9 0 R /Title (Page Two)
|
||||
>>
|
||||
endobj
|
||||
11 0 obj
|
||||
<<
|
||||
/Dest [ 5 0 R /Fit ] /Parent 8 0 R /Prev 10 0 R /Title (Page Three)
|
||||
>>
|
||||
endobj
|
||||
12 0 obj
|
||||
<<
|
||||
/Count 3 /Kids [ 3 0 R 4 0 R 5 0 R ] /Type /Pages
|
||||
>>
|
||||
endobj
|
||||
13 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 122
|
||||
>>
|
||||
stream
|
||||
Gap@Db6gL2'Lh3!@LZ0U8'7>U;'tH2cm;<+UO9dKg:K5pXY%ILno7bT=/&<K"<EU]9[SSm*P9LuAr1A`Y./=ub]S+JZn3-Xqn.)>^t)3an^%I4h`&g_Y!<hT~>endstream
|
||||
endobj
|
||||
14 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 124
|
||||
>>
|
||||
stream
|
||||
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CW4KISi<![7`#OB_qus.nXJpV`4oKb/`HKs]']P1$(("^Qh6`:R"4,>ElR/;4WeODY4S!3T'6Jc~>endstream
|
||||
endobj
|
||||
15 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 124
|
||||
>>
|
||||
stream
|
||||
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CW4KISi<![7`#OB_qus.nXJpV`4oKb/`HKs]']P1$(("^Qh6`:R"4@nhZ=/;4WeODY4S!3TQDK)~>endstream
|
||||
endobj
|
||||
xref
|
||||
0 16
|
||||
0000000000 65535 f
|
||||
0000000061 00000 n
|
||||
0000000092 00000 n
|
||||
0000000199 00000 n
|
||||
0000000394 00000 n
|
||||
0000000589 00000 n
|
||||
0000000784 00000 n
|
||||
0000000869 00000 n
|
||||
0000001130 00000 n
|
||||
0000001202 00000 n
|
||||
0000001289 00000 n
|
||||
0000001389 00000 n
|
||||
0000001479 00000 n
|
||||
0000001551 00000 n
|
||||
0000001764 00000 n
|
||||
0000001979 00000 n
|
||||
trailer
|
||||
<<
|
||||
/ID
|
||||
[<6132df6a3beacba675b566627d60fb2e><6132df6a3beacba675b566627d60fb2e>]
|
||||
% ReportLab generated PDF document -- digest (opensource)
|
||||
|
||||
/Info 7 0 R
|
||||
/Root 6 0 R
|
||||
/Size 16
|
||||
>>
|
||||
startxref
|
||||
2194
|
||||
%%EOF
|
||||
@@ -0,0 +1,106 @@
|
||||
%PDF-1.3
|
||||
%“Œ‹ž ReportLab Generated PDF document (opensource)
|
||||
1 0 obj
|
||||
<<
|
||||
/F1 2 0 R
|
||||
>>
|
||||
endobj
|
||||
2 0 obj
|
||||
<<
|
||||
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
|
||||
>>
|
||||
endobj
|
||||
3 0 obj
|
||||
<<
|
||||
/Contents 9 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
4 0 obj
|
||||
<<
|
||||
/Contents 10 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
5 0 obj
|
||||
<<
|
||||
/Contents 11 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
6 0 obj
|
||||
<<
|
||||
/PageMode /UseNone /Pages 8 0 R /Type /Catalog
|
||||
>>
|
||||
endobj
|
||||
7 0 obj
|
||||
<<
|
||||
/Author (anonymous) /CreationDate (D:20260917153218-04'00') /Creator (anonymous) /Keywords () /ModDate (D:20260917153218-04'00') /Producer (ReportLab PDF Library - \(opensource\))
|
||||
/Subject (unspecified) /Title (untitled) /Trapped /False
|
||||
>>
|
||||
endobj
|
||||
8 0 obj
|
||||
<<
|
||||
/Count 3 /Kids [ 3 0 R 4 0 R 5 0 R ] /Type /Pages
|
||||
>>
|
||||
endobj
|
||||
9 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
|
||||
>>
|
||||
stream
|
||||
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jt]&.8<uSu(.bn9(BF;Q*M*~>endstream
|
||||
endobj
|
||||
10 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
|
||||
>>
|
||||
stream
|
||||
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jruos8<uSu(.bn9(BF;_*M3~>endstream
|
||||
endobj
|
||||
11 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
|
||||
>>
|
||||
stream
|
||||
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?K!D1>8<uSu(.bn9(BF;m*M<~>endstream
|
||||
endobj
|
||||
xref
|
||||
0 12
|
||||
0000000000 65535 f
|
||||
0000000061 00000 n
|
||||
0000000092 00000 n
|
||||
0000000199 00000 n
|
||||
0000000392 00000 n
|
||||
0000000586 00000 n
|
||||
0000000780 00000 n
|
||||
0000000848 00000 n
|
||||
0000001109 00000 n
|
||||
0000001180 00000 n
|
||||
0000001405 00000 n
|
||||
0000001631 00000 n
|
||||
trailer
|
||||
<<
|
||||
/ID
|
||||
[<464fc7cfdf793a5b6d29e3d0d043c5a7><464fc7cfdf793a5b6d29e3d0d043c5a7>]
|
||||
% ReportLab generated PDF document -- digest (opensource)
|
||||
|
||||
/Info 7 0 R
|
||||
/Root 6 0 R
|
||||
/Size 12
|
||||
>>
|
||||
startxref
|
||||
1857
|
||||
%%EOF
|
||||
@@ -22,6 +22,23 @@ def _reset_tool_ratelimit():
|
||||
ratelimit.reset()
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _disable_web_cache():
|
||||
"""Web cache off by default: tests stay hermetic (no cross-test hits).
|
||||
|
||||
tests/test_web_cache.py re-enables it explicitly with a tmp path.
|
||||
"""
|
||||
from backend.tools import webcache
|
||||
|
||||
saved_path = os.environ.get("OBSIGATE_WEB_CACHE_PATH")
|
||||
os.environ["OBSIGATE_WEB_CACHE_TTL"] = "0"
|
||||
yield
|
||||
if saved_path is None:
|
||||
os.environ.pop("OBSIGATE_WEB_CACHE_PATH", None)
|
||||
else:
|
||||
os.environ["OBSIGATE_WEB_CACHE_PATH"] = saved_path
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clean_env():
|
||||
"""Ensure no vault env vars leak between tests — but preserve test vault config."""
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
/**
|
||||
* E2E tests for ObsiGate PDF inline viewer (BUG-060).
|
||||
*
|
||||
* Regression : la CSP posée par BUG-034 (`object-src 'none'`) bloque l'élément
|
||||
* `<embed>` qui servait le PDF. Le viewer doit rendre le stream dans une
|
||||
* `<iframe>` (autorisée par `frame-src 'self'`).
|
||||
*
|
||||
* Fixtures : `test_vault/sample-pdf.pdf` (3 pages, texte simple).
|
||||
*
|
||||
* Run (local):
|
||||
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js
|
||||
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js --headed
|
||||
*/
|
||||
|
||||
import { test, expect } from '@playwright/test';
|
||||
|
||||
const BASE = process.env.BASE_URL || 'http://localhost:2029';
|
||||
|
||||
const CREDS = {
|
||||
username: process.env.OBSIGATE_USER || 'admin',
|
||||
password: process.env.OBSIGATE_PASS || 'test123',
|
||||
};
|
||||
|
||||
async function login(page) {
|
||||
await page.goto(BASE);
|
||||
const loginForm = page.locator('#login-screen');
|
||||
await expect(loginForm).toBeVisible({ timeout: 5000 }).catch(() => {});
|
||||
if (await loginForm.isVisible()) {
|
||||
await page.fill('#login-username', CREDS.username);
|
||||
await page.fill('#login-password', CREDS.password);
|
||||
await page.click('#login-btn');
|
||||
}
|
||||
await page.waitForFunction(() => window.__OBSIGATE_BOOTED === true, { timeout: 20000 });
|
||||
}
|
||||
|
||||
async function openFile(page, vault, filePath) {
|
||||
const treeItem = page.locator(`.tree-item[data-vault="${vault}"][data-path="${filePath}"]`);
|
||||
// Le vault peut être replié (auth activée) : l'étendre avant de chercher le fichier.
|
||||
if (!(await treeItem.count())) {
|
||||
await page.locator(`.tree-item.vault-item[data-vault="${vault}"]`).first().click();
|
||||
await treeItem.waitFor({ state: 'attached', timeout: 8000 });
|
||||
}
|
||||
await treeItem.dblclick({ timeout: 5000 });
|
||||
await page.waitForTimeout(500);
|
||||
}
|
||||
|
||||
test.describe('PDF viewer — affichage inline (BUG-060)', () => {
|
||||
|
||||
test('ouvre un PDF dans une iframe (pas d\'<embed>) et le stream charge sans violation CSP', async ({ page }) => {
|
||||
const cspViolations = [];
|
||||
page.on('console', (msg) => {
|
||||
const text = msg.text();
|
||||
if (text.includes('Content Security Policy') && (text.includes('object-src') || text.includes('Refused'))) {
|
||||
cspViolations.push(text);
|
||||
}
|
||||
});
|
||||
|
||||
await login(page);
|
||||
|
||||
// Le stream est demandé au moment où le viewer monte l'iframe : enregistrer
|
||||
// l'écoute AVANT d'ouvrir le fichier (sinon la réponse est déjà passée).
|
||||
const streamResponsePromise = page.waitForResponse(
|
||||
(r) => r.url().includes('/pdf/stream') && (r.status() === 200 || r.status() === 206),
|
||||
{ timeout: 15000 },
|
||||
);
|
||||
|
||||
await openFile(page, 'TestVault', 'sample-pdf.pdf');
|
||||
|
||||
const iframe = page.locator('#content-area .pdf-iframe');
|
||||
await expect(iframe).toBeVisible({ timeout: 15000 });
|
||||
await expect(iframe).toHaveAttribute('src', /\/api\/file\/TestVault\/pdf\/stream\?path=/);
|
||||
|
||||
// La balise doit être une iframe (le <embed>/<object> serait bloqué par CSP)
|
||||
const tagName = await iframe.evaluate((el) => el.tagName);
|
||||
expect(tagName).toBe('IFRAME');
|
||||
expect(await page.locator('#content-area embed, #content-area object').count()).toBe(0);
|
||||
|
||||
// La barre d'outils indique le nombre de pages du PDF
|
||||
await expect(page.locator('.pdf-info')).toContainText('3 pages');
|
||||
|
||||
// Le stream est bien servi en application/pdf (200 ou 206 Range)
|
||||
const streamResp = await streamResponsePromise;
|
||||
expect(streamResp.headers()['content-type'] || '').toContain('application/pdf');
|
||||
|
||||
// Le cadre embarque réellement le document PDF (navigateur natif)
|
||||
const pdfFrame = await iframe.contentFrame();
|
||||
expect(pdfFrame).not.toBeNull();
|
||||
|
||||
// Aucune violation CSP liée à object-src pendant l'ouverture
|
||||
expect(cspViolations).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
test.describe('PDF viewer — TOC & plein largeur', () => {
|
||||
|
||||
test('les entrées de la TOC rechargent l\'iframe sur la page ciblée (#page=N)', async ({ page }) => {
|
||||
await login(page);
|
||||
await openFile(page, 'TestVault', 'sample-pdf-toc.pdf');
|
||||
|
||||
const tocLinks = page.locator('#content-area .pdf-toc a[data-page]');
|
||||
await expect(tocLinks).toHaveCount(3, { timeout: 10000 });
|
||||
|
||||
await page.locator('#content-area .pdf-toc a[data-page="3"]').click();
|
||||
// Le changement de query force un rechargement (un simple changement de
|
||||
// fragment est ignoré par le lecteur PDF natif), puis #page=3 est appliqué.
|
||||
await expect(page.locator('#content-area .pdf-iframe')).toHaveAttribute(
|
||||
'src',
|
||||
/\/pdf\/stream\?path=.*&_pdfpage=\d+#page=3$/,
|
||||
{ timeout: 5000 },
|
||||
);
|
||||
});
|
||||
|
||||
test('le PDF occupe toute la largeur quand la navigation est masquée', async ({ page }) => {
|
||||
await login(page);
|
||||
await openFile(page, 'TestVault', 'sample-pdf-toc.pdf');
|
||||
await expect(page.locator('#content-area .pdf-viewer-container')).toBeVisible({ timeout: 10000 });
|
||||
|
||||
// Masquer la barre de navigation (bouton réel).
|
||||
await page.locator('#sidebar-toggle-btn').click();
|
||||
await expect(page.locator('#sidebar')).toHaveClass(/hidden/);
|
||||
|
||||
const widths = await page.evaluate(() => {
|
||||
const area = document.getElementById('content-area');
|
||||
const cs = getComputedStyle(area);
|
||||
const container = document.querySelector('.pdf-viewer-container');
|
||||
const contentWidth =
|
||||
area.clientWidth - parseFloat(cs.paddingLeft) - parseFloat(cs.paddingRight);
|
||||
return {
|
||||
container: container.getBoundingClientRect().width,
|
||||
contentWidth,
|
||||
maxWidth: cs.maxWidth,
|
||||
};
|
||||
});
|
||||
|
||||
// Le plafond de lecture (1200px) ne doit plus s'appliquer au viewer PDF.
|
||||
expect(widths.maxWidth).toBe('none');
|
||||
expect(Math.abs(widths.container - widths.contentWidth)).toBeLessThan(2);
|
||||
});
|
||||
});
|
||||
@@ -302,6 +302,25 @@ async function main() {
|
||||
}
|
||||
});
|
||||
|
||||
await test("isScrollbarPress: content clicks never unpin, scrollbar drags do (BUG-059)", () => {
|
||||
const { isScrollbarPress } = bookslmMod;
|
||||
const container = document.createElement("div");
|
||||
container.getBoundingClientRect = () => ({
|
||||
left: 0, right: 800, top: 0, bottom: 600, width: 800, height: 600,
|
||||
});
|
||||
const link = document.createElement("a");
|
||||
// Click on a file link / steps toggle → must NOT unpin (the thread
|
||||
// would otherwise jump to the bottom when the padding is cleared).
|
||||
assert.equal(isScrollbarPress(link, 790, container), false);
|
||||
// Click on blank content area → must NOT unpin either.
|
||||
assert.equal(isScrollbarPress(container, 300, container), false);
|
||||
// Dragging the vertical scrollbar: target is the container itself and
|
||||
// the press lands in the right-edge gutter → unpin allowed.
|
||||
assert.equal(isScrollbarPress(container, 792, container), true);
|
||||
// Defensive: no container / no geometry → never unpin.
|
||||
assert.equal(isScrollbarPress(link, 790, null), false);
|
||||
});
|
||||
|
||||
await test("the placeholder render keeps the anchor (no scroll reset before first token)", async () => {
|
||||
// Regression: the assistant placeholder used to be rendered while
|
||||
// `_isLoading` was still false, so it took the "restore previous
|
||||
@@ -767,6 +786,29 @@ async function main() {
|
||||
panel.remove();
|
||||
});
|
||||
|
||||
// ── 10b. Fullscreen toggle must beat the persisted inline width ──
|
||||
await test("fullscreen button toggles the panel and CSS lifts the inline width", async () => {
|
||||
const { readFileSync } = await import("node:fs");
|
||||
const css = readFileSync(path.resolve(JS_DIR, "..", "style.css"), "utf-8");
|
||||
assert.match(
|
||||
css,
|
||||
/\.bookslm-panel\.fullscreen\s*\{\s*width:\s*100vw\s*!important;/,
|
||||
"the fullscreen width must override the inline width written by the resize handle",
|
||||
);
|
||||
|
||||
const b = new BooksLM();
|
||||
const panel = b._render();
|
||||
document.body.appendChild(panel);
|
||||
const btn = panel.querySelector(".bookslm-btn-fullscreen");
|
||||
btn.click();
|
||||
assert.ok(panel.classList.contains("fullscreen"), "fullscreen class added on first click");
|
||||
assert.equal(b._isFullscreen, true);
|
||||
btn.click();
|
||||
assert.ok(!panel.classList.contains("fullscreen"), "fullscreen class removed on second click");
|
||||
assert.equal(b._isFullscreen, false);
|
||||
panel.remove();
|
||||
});
|
||||
|
||||
// ── 11. Formatted markdown rendering ──
|
||||
await test("_renderMarkdown renders headings, lists, code and tables", () => {
|
||||
const b = new BooksLM();
|
||||
|
||||
@@ -380,6 +380,19 @@ test("style.css: #editor-body scrolls through the CodeMirror scroller only", ()
|
||||
"global .cm-scroller min-height override reintroduces the double scrollbar");
|
||||
});
|
||||
|
||||
// ── BUG-058: the line-number gutter follows the active theme ──
|
||||
test("style.css themes the CodeMirror line-number gutter with CSS variables", () => {
|
||||
const gutter = cssSrc.match(/\.cm-editor \.cm-gutters \{([^}]*)\}/);
|
||||
assert.ok(gutter, "themed .cm-gutters rule not found");
|
||||
assert.match(gutter[1], /background:\s*color-mix\([^;]*var\(--text-primary\)/,
|
||||
"gutter background must derive from the theme, not CodeMirror's hardcoded #f5f5f5");
|
||||
assert.match(gutter[1], /color:\s*var\(--text-secondary\)/);
|
||||
assert.match(gutter[1], /border-right:\s*1px solid var\(--border\)/);
|
||||
const active = cssSrc.match(/\.cm-editor \.cm-gutters \.cm-activeLineGutter \{([^}]*)\}/);
|
||||
assert.ok(active, "themed .cm-activeLineGutter rule not found");
|
||||
assert.match(active[1], /color-mix\([^;]*var\(--text-primary\)/);
|
||||
});
|
||||
|
||||
// ── BUG-054: the shared save button must not stay stuck on the spinner ──
|
||||
test("utils.js resetSaveButton restores the checkmark and re-enables the button", () => {
|
||||
const fn = utilsSrc.match(/function resetSaveButton\(\) \{([\s\S]*?)\n\}/);
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* ObsiGate — Viewer PDF non-regression tests (BUG-060).
|
||||
*
|
||||
* Static checks on the source of the PDF inline viewer:
|
||||
* - BUG-060 : the CSP set by BUG-034 puts `object-src 'none'`, which blocks
|
||||
* `<embed>`/`<object>`. The PDF body was therefore never rendered (blank
|
||||
* pages). The viewer must now use an `<iframe>` — permitted by
|
||||
* `frame-src 'self'` since the PDF stream URL is same-origin.
|
||||
*
|
||||
* Usage: node tests/frontend/pdf-viewer.test.mjs
|
||||
*/
|
||||
|
||||
import { strict as assert } from "node:assert";
|
||||
import { readFileSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
|
||||
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = path.join(__dirname, "..", "..");
|
||||
|
||||
const viewer = readFileSync(path.join(ROOT, "frontend", "js", "viewer.js"), "utf8");
|
||||
const main = readFileSync(path.join(ROOT, "backend", "main.py"), "utf8");
|
||||
const css = readFileSync(path.join(ROOT, "frontend", "style.css"), "utf8");
|
||||
|
||||
function test(label, fn) {
|
||||
try {
|
||||
fn();
|
||||
console.log(" \u2713 " + label);
|
||||
} catch (err) {
|
||||
console.error(" \u2717 " + label + "\n " + err.message);
|
||||
process.exitCode = 1;
|
||||
}
|
||||
}
|
||||
|
||||
// ── viewer.js : le rendu PDF passe par une iframe ─────────────────────────
|
||||
test("viewer.js — PDF branch renders the stream in an <iframe class=pdf-iframe>", () => {
|
||||
const block = viewer.match(/if \(data\.is_pdf\) \{([\s\S]*?)\n \}/);
|
||||
assert.ok(block, "PDF render block not found");
|
||||
assert.match(
|
||||
block[1],
|
||||
/<iframe src="\$\{pdfUrl\}" data-pdf-url="\$\{pdfUrl\}" class="pdf-iframe"/,
|
||||
"PDF must use <iframe>, not <embed>/<object> (CSP object-src 'none' otherwise blocks it)",
|
||||
);
|
||||
assert.match(
|
||||
block[1],
|
||||
/<iframe src="\$\{pdfUrl\}"/,
|
||||
"iframe src must come from the /pdf/stream URL",
|
||||
);
|
||||
});
|
||||
|
||||
test("viewer.js — no <embed>/<object> left in the source", () => {
|
||||
assert.doesNotMatch(viewer, /<embed\b/i, "<embed> is blocked by CSP object-src 'none'");
|
||||
assert.doesNotMatch(viewer, /<object\b/i, "<object> is blocked by CSP object-src 'none'");
|
||||
});
|
||||
|
||||
test("viewer.js — TOC links carry a data-page and never poke contentWindow", () => {
|
||||
assert.match(
|
||||
viewer,
|
||||
/<a href="#" data-page="\$\{item\.page\}">/,
|
||||
"TOC links must expose the target page via data-page",
|
||||
);
|
||||
assert.doesNotMatch(
|
||||
viewer,
|
||||
/contentWindow\.location\.hash\s*=/,
|
||||
"contentWindow is about:blank in the native PDF viewer: hash navigation never reaches the document",
|
||||
);
|
||||
});
|
||||
|
||||
test("viewer.js — navigatePdfToPage forces a reload with the #page fragment", () => {
|
||||
const fn = viewer.match(/export function navigatePdfToPage\(area, page\) \{([\s\S]*?)\n\}/);
|
||||
assert.ok(fn, "navigatePdfToPage helper not found");
|
||||
assert.match(fn[1], /data-pdf-url/, "the base URL must be preserved without the fragment");
|
||||
assert.match(
|
||||
fn[1],
|
||||
/_pdfpage=\$\{Date\.now\(\)\}#page=\$\{page\}/,
|
||||
"the query must change to force a reload (a fragment-only change is ignored by the native viewer)",
|
||||
);
|
||||
});
|
||||
|
||||
test("style.css — full-bleed viewers ignore the centered reading width", () => {
|
||||
assert.match(
|
||||
css,
|
||||
/\.sidebar\.hidden ~ \.content-wrapper \.content-area:has\(\.pdf-viewer-container\)/,
|
||||
"PDF viewer must fill the width when the navigation sidebar is hidden",
|
||||
);
|
||||
const rule = css.match(
|
||||
/\.content-area:has\(\.pdf-viewer-container\)[\s\S]*?\{([^}]*)\}/,
|
||||
);
|
||||
assert.ok(rule, "full-bleed rule not found");
|
||||
assert.match(rule[1], /max-width:\s*none/, "the 1200px reading cap must be lifted");
|
||||
});
|
||||
|
||||
// ── backend : la CSP autorise le cadre same-origin ─────────────────────────
|
||||
test("backend CSP — frame-src 'self' allows same-origin iframes", () => {
|
||||
const csp = main.match(/frame-src ([^";]+);/);
|
||||
assert.ok(csp, "CSP frame-src directive not found");
|
||||
assert.ok(
|
||||
csp[1].includes("'self'"),
|
||||
`frame-src must allow 'self' (found: ${csp[1]}) — otherwise the PDF iframe is blocked`,
|
||||
);
|
||||
});
|
||||
|
||||
test("backend CSP — object-src 'none' stays in place (no defusing)", () => {
|
||||
assert.match(
|
||||
main,
|
||||
/object-src 'none';/,
|
||||
"object-src must stay locked to 'none': the fix is moving to <iframe>, not weakening CSP",
|
||||
);
|
||||
});
|
||||
|
||||
// ── style.css : l'iframe garde une hauteur utile ───────────────────────────
|
||||
test("style.css — .pdf-iframe fills the viewer body", () => {
|
||||
const rule = css.match(/\.pdf-iframe\s*\{([^}]*)\}/);
|
||||
assert.ok(rule, ".pdf-iframe rule not found");
|
||||
assert.match(rule[1], /flex:\s*1/, "iframe must stretch to fill the available height");
|
||||
assert.match(rule[1], /min-height/, "iframe must keep its minimum height");
|
||||
});
|
||||
|
||||
if (process.exitCode) {
|
||||
console.error("\nPDF viewer tests FAILED");
|
||||
} else {
|
||||
console.log("\nAll PDF viewer tests passed.");
|
||||
}
|
||||
@@ -743,6 +743,33 @@ class TestSecretRedactor:
|
||||
result = redact_file_content("hello world this is safe")
|
||||
assert result == "hello world this is safe"
|
||||
|
||||
def test_git_sha_not_redacted(self):
|
||||
"""BUG-035: a bare git commit SHA must not be mangled."""
|
||||
from backend.secret_redactor import redact_file_content
|
||||
sha = "a1b2c3d4e5f60718293a4b5c6d7e8f9012345678"
|
||||
text = f"commit {sha}\nMerge: {sha}"
|
||||
assert redact_file_content(text) == text
|
||||
|
||||
def test_sha256_checksum_not_redacted(self):
|
||||
from backend.secret_redactor import redact_file_content
|
||||
digest = "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
|
||||
text = f"sha256:{digest} file.tar.gz"
|
||||
assert redact_file_content(text) == text
|
||||
|
||||
def test_hex_secret_in_context_is_redacted(self):
|
||||
from backend.secret_redactor import redact_file_content
|
||||
secret = "0123456789abcdef0123456789abcdef01234567"
|
||||
result = redact_file_content(f"api_key={secret}")
|
||||
assert secret not in result
|
||||
assert "MASQUÉ" in result
|
||||
|
||||
def test_ambiguous_bare_hex_left_intact(self):
|
||||
"""A 40-char hex with no secret/hash keyword stays untouched."""
|
||||
from backend.secret_redactor import redact_file_content
|
||||
blob = "deadbeefdeadbeefdeadbeefdeadbeefdeadbeef"
|
||||
text = f"value {blob} end"
|
||||
assert redact_file_content(text) == text
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# Static / PWA caching (Cloudflare / mobile freshness)
|
||||
|
||||
@@ -44,6 +44,22 @@ class TestPasswordHashing:
|
||||
result = hash_password("ab")
|
||||
assert result is not None
|
||||
|
||||
def test_argon2_memory_recalibrated(self):
|
||||
"""BUG-038: memory cost must stay at the OWASP 19 MiB recommendation."""
|
||||
from backend.auth.password import (
|
||||
ARGON2_MEMORY_COST_KIB,
|
||||
ARGON2_PARALLELISM,
|
||||
ARGON2_TIME_COST,
|
||||
ph,
|
||||
)
|
||||
|
||||
assert ARGON2_MEMORY_COST_KIB == 19456
|
||||
assert ARGON2_TIME_COST == 2
|
||||
assert ARGON2_PARALLELISM == 1
|
||||
assert ph.memory_cost == 19456
|
||||
assert ph.time_cost == 2
|
||||
assert ph.parallelism == 1
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# JWT Handler
|
||||
@@ -279,3 +295,61 @@ class TestMiddleware:
|
||||
assert check_vault_access("Vault1", user) is True
|
||||
assert check_vault_access("Vault3", user) is False
|
||||
assert check_vault_access("Vault1", nobody) is False
|
||||
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
# Insecure (auth-disabled) deployment guard — BUG-037
|
||||
# ═══════════════════════════════════════════════════════════════════
|
||||
|
||||
class TestInsecureAuthGuard:
|
||||
def test_is_loopback_host(self):
|
||||
from backend.auth.middleware import is_loopback_host
|
||||
|
||||
assert is_loopback_host(None) is True
|
||||
assert is_loopback_host("127.0.0.1") is True
|
||||
assert is_loopback_host("::1") is True
|
||||
assert is_loopback_host("[::1]") is True
|
||||
assert is_loopback_host("localhost") is True
|
||||
assert is_loopback_host("0.0.0.0") is False
|
||||
assert is_loopback_host("192.168.1.10") is False
|
||||
|
||||
def test_bind_host_from_argv(self):
|
||||
from backend.auth.middleware import bind_host_from_argv
|
||||
|
||||
assert bind_host_from_argv(
|
||||
["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8080"]
|
||||
) == "0.0.0.0"
|
||||
assert bind_host_from_argv(["uvicorn", "app", "--host=127.0.0.1"]) == "127.0.0.1"
|
||||
assert bind_host_from_argv(["uvicorn", "app"]) is None
|
||||
|
||||
def test_guard_refuses_public_bind_without_optin(self, monkeypatch):
|
||||
from backend import main
|
||||
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
|
||||
monkeypatch.delenv("OBSIGATE_ALLOW_INSECURE", raising=False)
|
||||
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
|
||||
with pytest.raises(RuntimeError):
|
||||
main._guard_insecure_auth()
|
||||
|
||||
def test_guard_allows_loopback(self, monkeypatch):
|
||||
from backend import main
|
||||
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
|
||||
monkeypatch.delenv("OBSIGATE_ALLOW_INSECURE", raising=False)
|
||||
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "127.0.0.1"])
|
||||
main._guard_insecure_auth() # must not raise
|
||||
|
||||
def test_guard_allows_explicit_optin(self, monkeypatch):
|
||||
from backend import main
|
||||
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
|
||||
monkeypatch.setenv("OBSIGATE_ALLOW_INSECURE", "true")
|
||||
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
|
||||
main._guard_insecure_auth() # must not raise
|
||||
|
||||
def test_guard_noop_when_auth_enabled(self, monkeypatch):
|
||||
from backend import main
|
||||
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
|
||||
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
|
||||
main._guard_insecure_auth() # must not raise
|
||||
|
||||
@@ -137,6 +137,42 @@ class TestLogin:
|
||||
})
|
||||
assert resp.status_code == 401
|
||||
|
||||
def test_locked_account_returns_401_not_429(self, auth_client, monkeypatch):
|
||||
"""BUG-039: a locked account must be indistinguishable from an unknown one."""
|
||||
import backend.auth.router as auth_router
|
||||
|
||||
monkeypatch.setattr(auth_router, "is_locked", lambda username: True)
|
||||
resp = auth_client.post("/api/auth/login", json={
|
||||
"username": "admin",
|
||||
"password": "chab30",
|
||||
})
|
||||
assert resp.status_code == 401
|
||||
assert "verrouill" not in resp.json()["detail"].lower()
|
||||
|
||||
def test_account_rate_limited_returns_401(self, auth_client, monkeypatch):
|
||||
"""BUG-039: per-account throttling must not reveal the account exists."""
|
||||
import backend.auth.router as auth_router
|
||||
|
||||
monkeypatch.setattr(auth_router, "is_account_rate_limited", lambda username: True)
|
||||
resp = auth_client.post("/api/auth/login", json={
|
||||
"username": "admin",
|
||||
"password": "chab30",
|
||||
})
|
||||
assert resp.status_code == 401
|
||||
|
||||
def test_inactive_account_returns_401(self, auth_client, monkeypatch):
|
||||
"""BUG-039: a disabled account answers like an unknown user."""
|
||||
import backend.auth.router as auth_router
|
||||
|
||||
monkeypatch.setattr(auth_router, "get_user", lambda username: {
|
||||
"username": username, "active": False, "password_hash": "x",
|
||||
})
|
||||
resp = auth_client.post("/api/auth/login", json={
|
||||
"username": "admin",
|
||||
"password": "chab30",
|
||||
})
|
||||
assert resp.status_code == 401
|
||||
|
||||
def test_login_remember_me(self, auth_client):
|
||||
resp = auth_client.post("/api/auth/login", json={
|
||||
"username": "admin",
|
||||
|
||||
@@ -73,6 +73,36 @@ def test_authenticate_websocket_invalid_token_returns_none(monkeypatch):
|
||||
assert authenticate_websocket(_StubWebSocket(cookies={"access_token": "garbage"})) is None
|
||||
|
||||
|
||||
def test_authenticate_websocket_query_token_rejected(monkeypatch):
|
||||
"""BUG-036: the access token must never be accepted from the query string."""
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
|
||||
from backend.auth.jwt_handler import create_access_token
|
||||
|
||||
token = create_access_token({
|
||||
"username": "u", "role": "user", "vaults": ["*"], "display_name": "U",
|
||||
})
|
||||
ws = _StubWebSocket(query={"token": token})
|
||||
assert authenticate_websocket(ws) is None
|
||||
|
||||
|
||||
def test_authenticate_websocket_cookie_token_accepted(monkeypatch):
|
||||
"""The HttpOnly access_token cookie remains the supported transport."""
|
||||
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
|
||||
import backend.auth.user_store as user_store
|
||||
from backend.auth.jwt_handler import create_access_token
|
||||
|
||||
token = create_access_token({
|
||||
"username": "u", "role": "user", "vaults": ["*"], "display_name": "U",
|
||||
})
|
||||
monkeypatch.setattr(user_store, "get_user", lambda username: {
|
||||
"username": username, "role": "user", "vaults": ["*"],
|
||||
"display_name": "U", "active": True,
|
||||
})
|
||||
user = authenticate_websocket(_StubWebSocket(cookies={"access_token": token}))
|
||||
assert user is not None
|
||||
assert user["username"] == "u"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Manager unit tests (no WebSocket transport)
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -127,6 +157,27 @@ async def test_on_message_rejects_oversized_update(tmp_path: Path):
|
||||
assert room.updates == []
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_on_message_rejects_oversized_raw(tmp_path: Path):
|
||||
"""BUG-036: oversized raw frames are dropped before parsing."""
|
||||
target = tmp_path / "note.md"
|
||||
target.write_text("x", encoding="utf-8")
|
||||
manager = _make_manager(tmp_path)
|
||||
room = CollabRoom(vault="V", path="note.md", file_path=target)
|
||||
client = _FakeClient(conn_id=1)
|
||||
|
||||
import backend.collab as collab_mod
|
||||
|
||||
original = collab_mod.MAX_MESSAGE_CHARS
|
||||
try:
|
||||
collab_mod.MAX_MESSAGE_CHARS = 10
|
||||
await manager._on_message(room, client, json.dumps({"type": "text", "text": "hello"}))
|
||||
finally:
|
||||
collab_mod.MAX_MESSAGE_CHARS = original
|
||||
|
||||
assert room.pending_text is None
|
||||
|
||||
|
||||
class _FakeWebSocket:
|
||||
def __init__(self):
|
||||
self.sent: list[dict] = []
|
||||
|
||||
@@ -0,0 +1,230 @@
|
||||
"""Unit tests for the connected sources (#92): Gitea & GitHub tools.
|
||||
|
||||
All HTTP calls are mocked (httpx.request monkeypatched) — the CI never talks
|
||||
to a real Gitea/GitHub instance.
|
||||
"""
|
||||
|
||||
import base64
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
|
||||
import backend.tools.connected as connected
|
||||
from backend.tools.context import ToolContext, ToolError, ToolMode
|
||||
from backend.tools.registry import get_tool
|
||||
|
||||
|
||||
class FakeResponse:
|
||||
def __init__(self, json_data: Any = None, status_code: int = 200):
|
||||
self._json = json_data
|
||||
self.status_code = status_code
|
||||
|
||||
def json(self):
|
||||
return self._json
|
||||
|
||||
def raise_for_status(self):
|
||||
if self.status_code >= 400:
|
||||
import httpx
|
||||
|
||||
raise httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
|
||||
|
||||
|
||||
def _ctx() -> ToolContext:
|
||||
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def gitea_env(monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_GITEA_URL", "https://git.example.net")
|
||||
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "tok-gitea")
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def github_env(monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "tok-gh")
|
||||
|
||||
|
||||
def _patch_request(monkeypatch, handler):
|
||||
def fake_request(method, url, headers=None, timeout=None, follow_redirects=False, **kw):
|
||||
captured = {"method": method, "url": str(url), "headers": headers or {},
|
||||
"params": kw.get("params")}
|
||||
return handler(captured)
|
||||
|
||||
monkeypatch.setattr(connected.httpx, "request", fake_request)
|
||||
|
||||
|
||||
class TestRegistration:
|
||||
@pytest.mark.parametrize("name", ["git_list_repos", "git_search_issues", "git_get_file"])
|
||||
def test_tools_registered_read(self, name):
|
||||
from backend.tools.context import ToolRisk
|
||||
|
||||
spec = get_tool(name)
|
||||
assert spec is not None
|
||||
assert spec.risk == ToolRisk.READ
|
||||
assert "gitea" in spec.description or "github" in spec.description.lower()
|
||||
|
||||
|
||||
class TestConfiguration:
|
||||
def test_gitea_requires_base_url(self, monkeypatch):
|
||||
monkeypatch.delenv("OBSIGATE_GITEA_URL", raising=False)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected._provider_base("gitea")
|
||||
assert ei.value.code == "provider_not_configured"
|
||||
|
||||
def test_unknown_provider_rejected(self):
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected._provider_base("gitlab")
|
||||
assert ei.value.code == "invalid_arguments"
|
||||
|
||||
def test_gitea_token_sent_as_header(self, gitea_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data={"data": []})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
|
||||
assert captured["headers"]["Authorization"] == "token tok-gitea"
|
||||
|
||||
|
||||
class TestListRepos:
|
||||
def test_gitea_search_endpoint(self, gitea_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data={"data": [
|
||||
{"name": "ObsiGate", "full_name": "bruno/ObsiGate",
|
||||
"html_url": "https://git.example.net/bruno/ObsiGate",
|
||||
"description": "vault gateway", "updated_at": "2026-09-01",
|
||||
"private": False},
|
||||
]})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
|
||||
assert "/api/v1/repos/search" in captured["url"]
|
||||
assert out["repos"][0]["name"] == "ObsiGate"
|
||||
assert out["count"] == 1
|
||||
|
||||
def test_github_single_repo(self, github_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data={
|
||||
"name": "ObsiGate", "full_name": "bruno/ObsiGate",
|
||||
"html_url": "https://github.com/bruno/ObsiGate",
|
||||
"description": "", "updated_at": "2026-09-02", "private": True,
|
||||
})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(
|
||||
provider="github", repo="bruno/ObsiGate"))
|
||||
assert captured["url"].endswith("/repos/bruno/ObsiGate")
|
||||
assert out["repos"][0]["full_name"] == "bruno/ObsiGate"
|
||||
assert out["repos"][0]["private"] is True
|
||||
|
||||
|
||||
class TestSearchIssues:
|
||||
def test_gitea_scoped_to_repo(self, gitea_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data=[
|
||||
{"number": 12, "title": "Bug affichage", "html_url": "https://x/12",
|
||||
"state": "open"},
|
||||
])
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
|
||||
provider="gitea", query="affichage", repo="bruno/ObsiGate"))
|
||||
assert "/repos/bruno/ObsiGate/issues" in captured["url"]
|
||||
assert out["issues"][0]["id"] == 12
|
||||
assert out["issues"][0]["pull_request"] is False
|
||||
|
||||
def test_github_search_syntax(self, github_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data={"items": [
|
||||
{"number": 5, "title": "Crash on save", "html_url": "https://gh/5",
|
||||
"state": "open", "pull_request": {"url": "x"}},
|
||||
]})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
|
||||
provider="github", query="crash", repo="bruno/ObsiGate", state="open"))
|
||||
assert "/search/issues" in captured["url"]
|
||||
assert "repo:bruno/ObsiGate" in captured["params"]["q"]
|
||||
assert out["issues"][0]["pull_request"] is True
|
||||
|
||||
|
||||
class TestGetFile:
|
||||
def test_gitea_base64_content_decoded(self, gitea_env, monkeypatch):
|
||||
payload = base64.b64encode("# Readme\n\nBonjour".encode()).decode()
|
||||
|
||||
def handler(_captured):
|
||||
return FakeResponse(json_data={
|
||||
"path": "README.md", "size": 15, "encoding": "base64", "content": payload,
|
||||
})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
out = connected.git_get_file(_ctx(), connected.GitGetFileInput(
|
||||
provider="gitea", repo="bruno/ObsiGate", path="README.md"))
|
||||
assert "Bonjour" in out["content"]
|
||||
assert out["truncated"] is False
|
||||
|
||||
def test_github_ref_parameter(self, github_env, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def handler(captured_req):
|
||||
captured.update(captured_req)
|
||||
return FakeResponse(json_data={
|
||||
"path": "a.md", "size": 1, "encoding": "base64",
|
||||
"content": base64.b64encode(b"x").decode(),
|
||||
})
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
connected.git_get_file(_ctx(), connected.GitGetFileInput(
|
||||
provider="github", repo="o/r", path="a.md", ref="v2.9.0"))
|
||||
assert captured["url"].endswith("?ref=v2.9.0")
|
||||
|
||||
def test_missing_repo_or_path_rejected(self, gitea_env):
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected.git_get_file(_ctx(), connected.GitGetFileInput(
|
||||
provider="gitea", repo="", path="a.md"))
|
||||
assert ei.value.code == "invalid_arguments"
|
||||
|
||||
|
||||
class TestErrors:
|
||||
def test_404_maps_to_not_found(self, gitea_env, monkeypatch):
|
||||
def handler(_captured):
|
||||
return FakeResponse(json_data={}, status_code=404)
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
|
||||
assert ei.value.code == "not_found"
|
||||
|
||||
def test_401_maps_to_permission_denied(self, gitea_env, monkeypatch):
|
||||
def handler(_captured):
|
||||
return FakeResponse(json_data={}, status_code=401)
|
||||
|
||||
_patch_request(monkeypatch, handler)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
|
||||
assert ei.value.code == "permission_denied"
|
||||
|
||||
def test_network_error_maps_to_tool_error(self, gitea_env, monkeypatch):
|
||||
import httpx
|
||||
|
||||
def fake_request(*a, **kw):
|
||||
raise httpx.ConnectError("down")
|
||||
|
||||
monkeypatch.setattr(connected.httpx, "request", fake_request)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
|
||||
assert ei.value.code == "connected_source_unavailable"
|
||||
@@ -0,0 +1,135 @@
|
||||
"""Unit tests for the bounded site crawler (#92): crawl_site (WRITE + confirmation)."""
|
||||
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
|
||||
import backend.tools.crawler as crawler
|
||||
from backend.tools.api import ToolConfirmationRequired, ToolContext, ToolError, call_tool
|
||||
|
||||
|
||||
class FakeResponse:
|
||||
def __init__(self, content: bytes = b"", status_code: int = 200,
|
||||
headers: dict | None = None):
|
||||
self.content = content
|
||||
self.status_code = status_code
|
||||
self.headers = headers or {"content-type": "text/html; charset=utf-8"}
|
||||
self.encoding = "utf-8"
|
||||
|
||||
def raise_for_status(self):
|
||||
pass
|
||||
|
||||
|
||||
def _ctx() -> ToolContext:
|
||||
return ToolContext(
|
||||
user={"username": "tester", "role": "admin", "vaults": ["*"]},
|
||||
audit_enabled=False,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def vault(tmp_path, monkeypatch):
|
||||
vault_dir = tmp_path / "Vault"
|
||||
vault_dir.mkdir()
|
||||
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
|
||||
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
|
||||
return vault_dir
|
||||
|
||||
|
||||
PAGE_A = (
|
||||
b"<html><head><title>Docs</title></head><body>"
|
||||
b"<p>Bienvenue sur la documentation.</p>"
|
||||
b'<a href="/page-b">Suite</a><a href="https://other.dev/x">ext</a>'
|
||||
b"</body></html>"
|
||||
)
|
||||
PAGE_B = (
|
||||
b"<html><head><title>Page B</title></head><body><p>Details ici.</p></body></html>"
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def local_urls(monkeypatch):
|
||||
"""Skip the DNS-based SSRF guard: test hosts are fake, HTTP is mocked."""
|
||||
monkeypatch.setattr(crawler, "_assert_public_http_url", lambda url: url)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def two_pages(monkeypatch, local_urls):
|
||||
def fake_get(url, **kw):
|
||||
url = str(url)
|
||||
if url.endswith("/page-b"):
|
||||
return FakeResponse(content=PAGE_B)
|
||||
return FakeResponse(content=PAGE_A)
|
||||
|
||||
monkeypatch.setattr(crawler.httpx, "get", fake_get)
|
||||
|
||||
|
||||
class TestConfirmation:
|
||||
def test_requires_confirmation(self, vault, two_pages):
|
||||
with pytest.raises(ToolConfirmationRequired):
|
||||
call_tool("crawl_site", _ctx(), {
|
||||
"url": "https://docs.example.dev/start",
|
||||
"vault": "Vault", "path": "Crawls/docs.md",
|
||||
})
|
||||
|
||||
|
||||
class TestCrawl:
|
||||
def test_saves_same_host_pages(self, vault, two_pages):
|
||||
out = call_tool("crawl_site", _ctx(), {
|
||||
"url": "https://docs.example.dev/start",
|
||||
"vault": "Vault", "path": "Crawls/docs.md",
|
||||
}, confirm=True)
|
||||
assert out.ok and out.data["pages"] == 2
|
||||
digest = (vault / "Crawls" / "docs.md").read_text(encoding="utf-8")
|
||||
assert "# Crawl de docs.example.dev" in digest
|
||||
assert "Bienvenue sur la documentation." in digest
|
||||
assert "Details ici." in digest
|
||||
assert "other.dev" not in digest
|
||||
|
||||
def test_max_pages_bound(self, vault, monkeypatch, local_urls):
|
||||
# A link farm: every page links to a new page — cap at max_pages.
|
||||
def fake_get(url, **kw):
|
||||
url = str(url)
|
||||
n = int(url.rsplit("/", 1)[-1] or 0)
|
||||
return FakeResponse(
|
||||
content=f"<html><head><title>P{n}</title></head><body>"
|
||||
f"<p>page {n}</p><a href=\"/{n + 1}\">next</a></body></html>".encode())
|
||||
|
||||
monkeypatch.setattr(crawler.httpx, "get", fake_get)
|
||||
out = call_tool("crawl_site", _ctx(), {
|
||||
"url": "https://farm.example.dev/0",
|
||||
"vault": "Vault", "path": "farm.md", "max_pages": 3,
|
||||
}, confirm=True)
|
||||
assert out.ok and out.data["pages"] == 3
|
||||
|
||||
def test_no_pages_recovered(self, vault, monkeypatch, local_urls):
|
||||
def dead_get(*a, **kw):
|
||||
raise crawler.httpx.ConnectError("down")
|
||||
|
||||
monkeypatch.setattr(crawler.httpx, "get", dead_get)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
call_tool("crawl_site", _ctx(), {
|
||||
"url": "https://dead.example.dev/", "vault": "Vault", "path": "x.md",
|
||||
}, confirm=True)
|
||||
assert ei.value.code == "crawl_failed"
|
||||
|
||||
def test_internal_url_rejected(self, vault):
|
||||
with pytest.raises(ToolError) as ei:
|
||||
call_tool("crawl_site", _ctx(), {
|
||||
"url": "http://127.0.0.1:8080/", "vault": "Vault", "path": "x.md",
|
||||
}, confirm=True)
|
||||
assert ei.value.code in ("ssrf_blocked", "dns_error")
|
||||
|
||||
def test_binary_content_skipped(self, vault, monkeypatch, local_urls):
|
||||
def fake_get(url, **kw):
|
||||
url = str(url)
|
||||
if url.endswith("/x.pdf"):
|
||||
return FakeResponse(content=b"%PDF-1.4", headers={"content-type": "application/pdf"})
|
||||
return FakeResponse(content=PAGE_A)
|
||||
|
||||
monkeypatch.setattr(crawler.httpx, "get", fake_get)
|
||||
out = call_tool("crawl_site", _ctx(), {
|
||||
"url": "https://docs.example.dev/start",
|
||||
"vault": "Vault", "path": "docs.md", "max_pages": 5,
|
||||
}, confirm=True)
|
||||
assert out.ok and out.data["pages"] >= 1
|
||||
@@ -0,0 +1,197 @@
|
||||
"""Unit tests for the document-production tools (#92): create_xlsx, create_docx,
|
||||
create_csv, create_pdf — WRITE risk, confirmation gating, vault persistence."""
|
||||
|
||||
import csv
|
||||
import io
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from backend.tools.api import (
|
||||
ToolConfirmationRequired,
|
||||
ToolContext,
|
||||
ToolError,
|
||||
call_tool,
|
||||
get_tool,
|
||||
)
|
||||
from backend.tools.context import ToolRisk
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def vault(tmp_path, monkeypatch):
|
||||
"""A minimal configured vault (index entry patched, no full build)."""
|
||||
vault_dir = tmp_path / "Vault"
|
||||
vault_dir.mkdir()
|
||||
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
|
||||
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
|
||||
return vault_dir
|
||||
|
||||
|
||||
def _ctx() -> ToolContext:
|
||||
return ToolContext(
|
||||
user={"username": "tester", "role": "admin", "vaults": ["*"]},
|
||||
audit_enabled=False,
|
||||
)
|
||||
|
||||
|
||||
class TestRegistry:
|
||||
@pytest.mark.parametrize("name", ["create_xlsx", "create_docx", "create_csv", "create_pdf"])
|
||||
def test_write_risk_and_confirmation(self, name):
|
||||
spec = get_tool(name)
|
||||
assert spec is not None
|
||||
assert spec.risk == ToolRisk.WRITE
|
||||
assert spec.requires_confirmation is True
|
||||
|
||||
|
||||
class TestConfirmationGating:
|
||||
def test_csv_requires_confirmation(self, vault):
|
||||
with pytest.raises(ToolConfirmationRequired):
|
||||
call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "data.csv",
|
||||
"rows": [["a", "b"], [1, 2]],
|
||||
})
|
||||
|
||||
def test_pdf_requires_confirmation(self, vault):
|
||||
with pytest.raises(ToolConfirmationRequired):
|
||||
call_tool("create_pdf", _ctx(), {
|
||||
"vault": "Vault", "path": "doc.pdf", "title": "T", "content": "# H\npara",
|
||||
})
|
||||
|
||||
|
||||
class TestCreateCsv:
|
||||
def test_creates_file_in_vault(self, vault):
|
||||
out = call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "Exports/data.csv",
|
||||
"rows": [["nom", "score"], ["alice", 12], ["bob", 9.5]],
|
||||
}, confirm=True)
|
||||
assert out.ok and out.data["success"] is True
|
||||
path = vault / "Exports" / "data.csv"
|
||||
assert path.exists()
|
||||
rows = list(csv.reader(io.StringIO(path.read_text(encoding="utf-8"))))
|
||||
assert rows[0] == ["nom", "score"]
|
||||
assert rows[2] == ["bob", "9.5"]
|
||||
|
||||
def test_semicolon_delimiter(self, vault):
|
||||
call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "d.csv", "delimiter": ";",
|
||||
"rows": [["a", "b"], [1, 2]],
|
||||
}, confirm=True)
|
||||
content = (vault / "d.csv").read_text(encoding="utf-8")
|
||||
assert "a;b" in content
|
||||
|
||||
def test_wrong_extension_rejected(self, vault):
|
||||
with pytest.raises(ToolError) as ei:
|
||||
call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "d.txt", "rows": [["a"], [1]],
|
||||
}, confirm=True)
|
||||
assert ei.value.code == "invalid_arguments"
|
||||
|
||||
def test_empty_rows_rejected(self, vault):
|
||||
with pytest.raises(ToolError):
|
||||
call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "d.csv", "rows": [],
|
||||
}, confirm=True)
|
||||
|
||||
|
||||
class TestCreateXlsx:
|
||||
def test_creates_readable_workbook(self, vault):
|
||||
call_tool("create_xlsx", _ctx(), {
|
||||
"vault": "Vault", "path": "Rapports/budget.xlsx",
|
||||
"rows": [["item", "cout"], ["serveur", 1200], ["licence", 300]],
|
||||
"sheet_name": "Budget",
|
||||
}, confirm=True)
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(vault / "Rapports" / "budget.xlsx")
|
||||
ws = wb.active
|
||||
assert ws.title == "Budget"
|
||||
assert ws.cell(row=1, column=1).value == "item"
|
||||
assert ws.cell(row=2, column=2).value == 1200
|
||||
|
||||
def test_wrong_extension_rejected(self, vault):
|
||||
with pytest.raises(ToolError):
|
||||
call_tool("create_xlsx", _ctx(), {
|
||||
"vault": "Vault", "path": "b.docx", "rows": [["a"], [1]],
|
||||
}, confirm=True)
|
||||
|
||||
|
||||
class TestCreateDocx:
|
||||
def test_creates_readable_document(self, vault):
|
||||
call_tool("create_docx", _ctx(), {
|
||||
"vault": "Vault", "path": "rapport.docx",
|
||||
"title": "Rapport hebdo", "paragraphs": ["Premier point.", "Second point."],
|
||||
}, confirm=True)
|
||||
import docx as docx_lib
|
||||
|
||||
doc = docx_lib.Document(str(vault / "rapport.docx"))
|
||||
texts = [p.text for p in doc.paragraphs]
|
||||
assert "Rapport hebdo" in texts
|
||||
assert "Second point." in texts
|
||||
|
||||
def test_no_paragraphs_rejected(self, vault):
|
||||
with pytest.raises(ToolError):
|
||||
call_tool("create_docx", _ctx(), {
|
||||
"vault": "Vault", "path": "r.docx", "paragraphs": [],
|
||||
}, confirm=True)
|
||||
|
||||
|
||||
class TestCreatePdf:
|
||||
def test_creates_valid_pdf(self, vault):
|
||||
call_tool("create_pdf", _ctx(), {
|
||||
"vault": "Vault", "path": "docs/archi.pdf",
|
||||
"title": "Architecture", "content": "# Titre\n\nUn paragraphe.\n## Sous-titre\nAutre texte.",
|
||||
}, confirm=True)
|
||||
raw = (vault / "docs" / "archi.pdf").read_bytes()
|
||||
assert raw.startswith(b"%PDF")
|
||||
|
||||
def test_long_content_truncated(self, vault):
|
||||
call_tool("create_pdf", _ctx(), {
|
||||
"vault": "Vault", "path": "big.pdf", "title": "T", "content": "x" * 500_000,
|
||||
}, confirm=True)
|
||||
assert (vault / "big.pdf").exists()
|
||||
|
||||
def test_markdown_tables_go_through_the_export_pipeline(self, vault, monkeypatch):
|
||||
# The document-page pipeline (mistune tables + WeasyPrint) must be
|
||||
# used when available: capture the HTML handed to the PDF generator.
|
||||
import sys
|
||||
import types
|
||||
|
||||
captured = {}
|
||||
fake = types.ModuleType("backend.pdf_export")
|
||||
|
||||
def fake_build(html, title, **kw):
|
||||
captured["html"] = html
|
||||
captured["title"] = title
|
||||
return "<html>" + html + "</html>"
|
||||
|
||||
fake.build_pdf_html = fake_build
|
||||
fake.generate_pdf = lambda html, title=None, **kw: b"%PDF-fake"
|
||||
monkeypatch.setitem(sys.modules, "backend.pdf_export", fake)
|
||||
out = call_tool("create_pdf", _ctx(), {
|
||||
"vault": "Vault", "path": "table.pdf", "title": "Rapport",
|
||||
"content": "# T\n\n| a | b |\n|---|---|\n| 1 | 2 |",
|
||||
}, confirm=True)
|
||||
assert out.ok
|
||||
assert "<table>" in captured["html"]
|
||||
assert captured["title"] == "Rapport"
|
||||
assert (vault / "table.pdf").read_bytes() == b"%PDF-fake"
|
||||
|
||||
def test_reportlab_fallback_when_weasyprint_missing(self, vault, monkeypatch):
|
||||
# sys.modules[name] = None makes `from backend.pdf_export import …`
|
||||
# raise ImportError → the simplified renderer must take over.
|
||||
import sys
|
||||
|
||||
monkeypatch.setitem(sys.modules, "backend.pdf_export", None)
|
||||
call_tool("create_pdf", _ctx(), {
|
||||
"vault": "Vault", "path": "fb.pdf", "title": "T", "content": "# H\ntexte",
|
||||
}, confirm=True)
|
||||
assert (vault / "fb.pdf").read_bytes().startswith(b"%PDF")
|
||||
|
||||
|
||||
class TestVaultSafety:
|
||||
def test_path_outside_vault_rejected(self, vault):
|
||||
with pytest.raises(ToolError) as ei:
|
||||
call_tool("create_csv", _ctx(), {
|
||||
"vault": "Vault", "path": "../outside.csv", "rows": [["a"], [1]],
|
||||
}, confirm=True)
|
||||
assert ei.value.code in ("path_outside_vault", "invalid_arguments", "tool_execution_error")
|
||||
+90
-3
@@ -275,6 +275,7 @@ class TestPdfIndexing:
|
||||
"""When a vault directory is scanned with .pdf files, they appear in files list.
|
||||
|
||||
Uses the public _scan_vault() helper directly — no global state needed.
|
||||
BUG-040: text extraction is deferred, so the scan only carries metadata.
|
||||
"""
|
||||
from backend.indexer import _scan_vault
|
||||
|
||||
@@ -286,10 +287,96 @@ class TestPdfIndexing:
|
||||
names = {f["path"] for f in result["files"]}
|
||||
assert "a.pdf" in names
|
||||
assert "b.pdf" in names
|
||||
# The PDF content should have been extracted.
|
||||
# The scan defers the expensive text extraction.
|
||||
a_file = next(f for f in result["files"] if f["path"] == "a.pdf")
|
||||
assert "ObsiGate test PDF" in (a_file.get("content") or "")
|
||||
assert "uniqueword0" in (a_file.get("content") or "")
|
||||
assert a_file["content"] == ""
|
||||
assert a_file["pdf_text_pending"] is True
|
||||
|
||||
|
||||
class TestPdfLazyEnrichment:
|
||||
"""BUG-040: PDF text is extracted in a deferred background pass."""
|
||||
|
||||
def test_scan_defers_pdf_text_extraction(self, pdf_dir: Path, tmp_path: Path):
|
||||
from backend.indexer import _scan_vault
|
||||
|
||||
vault_root = tmp_path / "vault"
|
||||
vault_root.mkdir()
|
||||
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
|
||||
result = _scan_vault("v", str(vault_root), {})
|
||||
a_file = next(f for f in result["files"] if f["path"] == "a.pdf")
|
||||
assert a_file["content"] == ""
|
||||
assert a_file["content_preview"] == ""
|
||||
assert a_file["pdf_text_pending"] is True
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_enrich_pdf_texts_fills_content_and_clears_flag(
|
||||
self, pdf_dir: Path, tmp_path: Path
|
||||
):
|
||||
import backend.indexer as idx
|
||||
|
||||
vault_root = tmp_path / "vault"
|
||||
vault_root.mkdir()
|
||||
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
|
||||
file_info = {
|
||||
"path": "a.pdf",
|
||||
"title": "a",
|
||||
"tags": [],
|
||||
"content": "",
|
||||
"content_preview": "",
|
||||
"size": 0,
|
||||
"modified": "",
|
||||
"extension": ".pdf",
|
||||
"pdf_text_pending": True,
|
||||
}
|
||||
with idx._index_lock:
|
||||
idx.index["LazyV"] = {
|
||||
"files": [file_info],
|
||||
"tags": {},
|
||||
"path": str(vault_root),
|
||||
"paths": [],
|
||||
}
|
||||
try:
|
||||
count = await idx.enrich_pdf_texts("LazyV")
|
||||
assert count == 1
|
||||
assert "ObsiGate test PDF" in file_info["content"]
|
||||
assert "uniqueword0" in file_info["content"]
|
||||
assert file_info["content_preview"]
|
||||
assert "pdf_text_pending" not in file_info
|
||||
finally:
|
||||
with idx._index_lock:
|
||||
idx.index.pop("LazyV", None)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_enrich_pdf_texts_skips_other_vaults(self, pdf_dir: Path, tmp_path: Path):
|
||||
import backend.indexer as idx
|
||||
|
||||
vault_root = tmp_path / "vault"
|
||||
vault_root.mkdir()
|
||||
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
|
||||
file_info = {
|
||||
"path": "a.pdf",
|
||||
"title": "a",
|
||||
"tags": [],
|
||||
"content": "",
|
||||
"content_preview": "",
|
||||
"size": 0,
|
||||
"modified": "",
|
||||
"extension": ".pdf",
|
||||
"pdf_text_pending": True,
|
||||
}
|
||||
with idx._index_lock:
|
||||
idx.index["OtherV"] = {
|
||||
"files": [file_info],
|
||||
"tags": {},
|
||||
"path": str(vault_root),
|
||||
"paths": [],
|
||||
}
|
||||
try:
|
||||
assert await idx.enrich_pdf_texts("LazyV") == 0
|
||||
assert file_info["content"] == ""
|
||||
finally:
|
||||
with idx._index_lock:
|
||||
idx.index.pop("OtherV", None)
|
||||
|
||||
|
||||
# ── Search filter `ext:` ───────────────────────────────────────────────────
|
||||
|
||||
@@ -0,0 +1,163 @@
|
||||
"""Unit tests for the tool/connected-source key store (#103).
|
||||
|
||||
Covers ``backend.tools.secrets`` (precedence, masking, whitelist) and the
|
||||
``/api/config/tool-keys`` endpoints (masked GET, POST, DELETE, admin-only).
|
||||
"""
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from backend.tools.secrets import (
|
||||
TOOL_KEY_NAMES,
|
||||
delete_tool_key,
|
||||
get_tool_key,
|
||||
is_secret_name,
|
||||
mask_value,
|
||||
set_tool_key,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def key_store(tmp_path, monkeypatch):
|
||||
"""Isolated key store directory."""
|
||||
monkeypatch.setenv("OBSIGATE_DATA_DIR", str(tmp_path))
|
||||
yield tmp_path
|
||||
|
||||
|
||||
def _write_store(tmp_path, data):
|
||||
(tmp_path / "api_keys.json").write_text(json.dumps(data), encoding="utf-8")
|
||||
|
||||
|
||||
class TestGetToolKey:
|
||||
def test_stored_value_takes_precedence_over_env(self, key_store, monkeypatch):
|
||||
_write_store(key_store, {"OBSIGATE_GITHUB_TOKEN": "stored-token"})
|
||||
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "env-token")
|
||||
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == "stored-token"
|
||||
|
||||
def test_env_fallback_when_not_stored(self, key_store, monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "env-token")
|
||||
assert get_tool_key("OBSIGATE_GITEA_TOKEN") == "env-token"
|
||||
|
||||
def test_missing_everywhere_returns_empty(self, key_store, monkeypatch):
|
||||
monkeypatch.delenv("OBSIGATE_GITHUB_TOKEN", raising=False)
|
||||
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == ""
|
||||
|
||||
def test_non_whitelisted_name_uses_env_only(self, key_store, monkeypatch):
|
||||
_write_store(key_store, {"OTHER_KEY": "stored"})
|
||||
monkeypatch.setenv("OTHER_KEY", "env")
|
||||
assert get_tool_key("OTHER_KEY") == "env"
|
||||
|
||||
def test_whitelist_covers_expected_names(self):
|
||||
assert set(TOOL_KEY_NAMES) == {
|
||||
"OBSIGATE_TAVILY_API_KEY",
|
||||
"OBSIGATE_BRAVE_API_KEY",
|
||||
"OBSIGATE_SERPAPI_API_KEY",
|
||||
"OBSIGATE_EXA_API_KEY",
|
||||
"OBSIGATE_GITEA_URL",
|
||||
"OBSIGATE_GITEA_TOKEN",
|
||||
"OBSIGATE_GITHUB_TOKEN",
|
||||
}
|
||||
|
||||
|
||||
class TestSetDelete:
|
||||
def test_set_then_get_roundtrip(self, key_store):
|
||||
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
|
||||
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == "tvly-1234"
|
||||
|
||||
def test_set_empty_value_deletes_entry(self, key_store):
|
||||
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
|
||||
set_tool_key("OBSIGATE_TAVILY_API_KEY", "")
|
||||
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == ""
|
||||
assert "OBSIGATE_TAVILY_API_KEY" not in json.loads(
|
||||
(key_store / "api_keys.json").read_text(encoding="utf-8")
|
||||
)
|
||||
|
||||
def test_delete_removes_and_reports(self, key_store):
|
||||
set_tool_key("OBSIGATE_GITEA_URL", "https://git.example.net")
|
||||
assert delete_tool_key("OBSIGATE_GITEA_URL") is True
|
||||
assert delete_tool_key("OBSIGATE_GITEA_URL") is False
|
||||
|
||||
def test_unknown_name_rejected(self, key_store):
|
||||
with pytest.raises(ValueError):
|
||||
set_tool_key("NOT_WHITELISTED", "x")
|
||||
with pytest.raises(ValueError):
|
||||
delete_tool_key("NOT_WHITELISTED")
|
||||
|
||||
def test_store_keeps_other_entries(self, key_store):
|
||||
_write_store(key_store, {"DEEPSEEK_API_KEY": "sk-existing"})
|
||||
set_tool_key("OBSIGATE_EXA_API_KEY", "exa-key")
|
||||
data = json.loads((key_store / "api_keys.json").read_text(encoding="utf-8"))
|
||||
assert data["DEEPSEEK_API_KEY"] == "sk-existing"
|
||||
assert data["OBSIGATE_EXA_API_KEY"] == "exa-key"
|
||||
|
||||
|
||||
class TestMasking:
|
||||
def test_urls_returned_clear(self):
|
||||
assert mask_value("OBSIGATE_GITEA_URL", "https://git.example.net") == \
|
||||
"https://git.example.net"
|
||||
|
||||
def test_tokens_masked(self):
|
||||
masked = mask_value("OBSIGATE_GITHUB_TOKEN", "ghp_abcdefgh1234")
|
||||
assert masked.startswith("ghp_")
|
||||
assert "abcdefgh1234" not in masked
|
||||
assert "..." in masked
|
||||
|
||||
def test_short_secret_fully_masked(self):
|
||||
assert mask_value("OBSIGATE_EXA_API_KEY", "abc") == "***"
|
||||
|
||||
def test_secret_detection(self):
|
||||
assert is_secret_name("OBSIGATE_GITEA_TOKEN")
|
||||
assert is_secret_name("OBSIGATE_TAVILY_API_KEY")
|
||||
assert not is_secret_name("OBSIGATE_GITEA_URL")
|
||||
|
||||
|
||||
class TestToolKeysAPI:
|
||||
def _login(self, admin_client):
|
||||
resp = admin_client.post(
|
||||
"/api/auth/login", json={"username": "admin", "password": "chab30"}
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
token = resp.json()["access_token"]
|
||||
return {"Authorization": f"Bearer {token}"}
|
||||
|
||||
def test_get_masks_tokens_shows_urls(self, admin_client, key_store):
|
||||
headers = self._login(admin_client)
|
||||
_write_store(key_store, {
|
||||
"OBSIGATE_GITEA_URL": "https://git.example.net",
|
||||
"OBSIGATE_GITHUB_TOKEN": "ghp_abcdefgh1234",
|
||||
})
|
||||
resp = admin_client.get("/api/config/tool-keys", headers=headers)
|
||||
assert resp.status_code == 200
|
||||
data = resp.json()
|
||||
assert data["OBSIGATE_GITEA_URL"] == "https://git.example.net"
|
||||
assert "abcdefgh1234" not in data["OBSIGATE_GITHUB_TOKEN"]
|
||||
assert data["OBSIGATE_GITHUB_TOKEN"].startswith("ghp_")
|
||||
|
||||
def test_post_roundtrip_then_delete(self, admin_client, key_store):
|
||||
headers = self._login(admin_client)
|
||||
resp = admin_client.post(
|
||||
"/api/config/tool-keys",
|
||||
headers=headers,
|
||||
json={"OBSIGATE_EXA_API_KEY": "exa-key-1234"},
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
assert resp.json()["status"] == "ok"
|
||||
assert get_tool_key("OBSIGATE_EXA_API_KEY") == "exa-key-1234"
|
||||
|
||||
resp = admin_client.delete("/api/config/tool-keys/OBSIGATE_EXA_API_KEY", headers=headers)
|
||||
assert resp.status_code == 200
|
||||
assert resp.json()["status"] == "deleted"
|
||||
assert get_tool_key("OBSIGATE_EXA_API_KEY") == ""
|
||||
|
||||
def test_post_unknown_name_rejected(self, admin_client, key_store):
|
||||
headers = self._login(admin_client)
|
||||
resp = admin_client.post(
|
||||
"/api/config/tool-keys", headers=headers, json={"MY_SECRET": "x"}
|
||||
)
|
||||
assert resp.status_code == 400
|
||||
|
||||
def test_requires_admin(self, admin_client, key_store):
|
||||
resp = admin_client.get("/api/config/tool-keys", headers={})
|
||||
assert resp.status_code in (401, 403)
|
||||
@@ -0,0 +1,82 @@
|
||||
"""Unit tests for the SQLite web cache (backend.tools.webcache, #92)."""
|
||||
|
||||
import pytest
|
||||
|
||||
from backend.tools import webcache
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def cache_enabled(tmp_path, monkeypatch):
|
||||
"""Enable the cache against an isolated file with a short TTL."""
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
|
||||
webcache._schema_ready = False
|
||||
yield
|
||||
webcache._schema_ready = False
|
||||
|
||||
|
||||
class TestCacheKey:
|
||||
def test_deterministic_and_payload_sensitive(self):
|
||||
a = webcache.cache_key("search", {"q": "pizza", "page": 1})
|
||||
b = webcache.cache_key("search", {"page": 1, "q": "pizza"})
|
||||
c = webcache.cache_key("search", {"q": "pasta", "page": 1})
|
||||
d = webcache.cache_key("fetch", {"q": "pizza", "page": 1})
|
||||
assert a == b
|
||||
assert a != c
|
||||
assert a != d
|
||||
|
||||
|
||||
class TestCacheRoundTrip:
|
||||
def test_set_get_roundtrip(self, cache_enabled):
|
||||
key = webcache.cache_key("search", {"q": "x"})
|
||||
webcache.cache_set(key, {"results": [1, 2, 3], "provider": "tavily"})
|
||||
assert webcache.cache_get(key) == {"results": [1, 2, 3], "provider": "tavily"}
|
||||
|
||||
def test_miss_returns_none(self, cache_enabled):
|
||||
assert webcache.cache_get("search:unknown") is None
|
||||
|
||||
def test_overwrite_updates_value(self, cache_enabled):
|
||||
key = webcache.cache_key("search", {"q": "x"})
|
||||
webcache.cache_set(key, {"v": 1})
|
||||
webcache.cache_set(key, {"v": 2})
|
||||
assert webcache.cache_get(key) == {"v": 2}
|
||||
|
||||
def test_ttl_expiry(self, tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
|
||||
webcache._schema_ready = False
|
||||
key = webcache.cache_key("search", {"q": "x"})
|
||||
webcache.cache_set(key, {"v": 1})
|
||||
assert webcache.cache_get(key) == {"v": 1}
|
||||
import time
|
||||
|
||||
time.sleep(0.15)
|
||||
assert webcache.cache_get(key) is None
|
||||
|
||||
def test_disabled_when_ttl_zero(self, cache_enabled, monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0")
|
||||
key = webcache.cache_key("search", {"q": "x"})
|
||||
webcache.cache_set(key, {"v": 1})
|
||||
assert webcache.cache_get(key) is None
|
||||
|
||||
def test_purge_expired(self, cache_enabled, monkeypatch):
|
||||
key = webcache.cache_key("search", {"q": "x"})
|
||||
webcache.cache_set(key, {"v": 1})
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
|
||||
import time
|
||||
|
||||
time.sleep(0.15)
|
||||
assert webcache.purge_expired() >= 1
|
||||
assert webcache.cache_get(key) is None
|
||||
|
||||
def test_corrupt_db_degrades_silently(self, tmp_path, monkeypatch):
|
||||
# A directory as cache file breaks sqlite3.connect: the cache must
|
||||
# disable itself instead of breaking the tools.
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path))
|
||||
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
|
||||
webcache._schema_ready = False
|
||||
try:
|
||||
assert webcache.cache_get("search:x") is None
|
||||
webcache.cache_set("search:x", {"v": 1})
|
||||
finally:
|
||||
webcache._schema_ready = False
|
||||
@@ -0,0 +1,155 @@
|
||||
"""Unit tests for the keyed web-search providers (#92): Tavily, Brave,
|
||||
SerpAPI, Exa — plus provider ordering and the transient retry."""
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
|
||||
import backend.tools.web as web
|
||||
from backend.tools.context import ToolContext, ToolError, ToolMode
|
||||
|
||||
|
||||
class FakeResponse:
|
||||
def __init__(self, json_data=None, status_code=200, content=b""):
|
||||
self._json = json_data
|
||||
self.status_code = status_code
|
||||
self.content = content
|
||||
self.encoding = "utf-8"
|
||||
self.headers = {}
|
||||
|
||||
def json(self):
|
||||
return self._json
|
||||
|
||||
def raise_for_status(self):
|
||||
if self.status_code >= 400:
|
||||
raise web.httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
|
||||
|
||||
|
||||
def _ctx() -> ToolContext:
|
||||
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
|
||||
|
||||
|
||||
def _no_fallback(monkeypatch):
|
||||
"""Limit the chain to the provider under test (no searxng/ddg/bing noise)."""
|
||||
monkeypatch.setattr(web, "WEB_FALLBACK_ENABLED", False)
|
||||
monkeypatch.setattr(web, "SEARXNG_URL", "http://searxng.invalid")
|
||||
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: (_ for _ in ()).throw(
|
||||
httpx.ConnectError("offline")))
|
||||
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: (_ for _ in ()).throw(
|
||||
httpx.ConnectError("offline")))
|
||||
|
||||
|
||||
class TestKeyedProviderParsers:
|
||||
def test_tavily_maps_results(self, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def fake_post(url, json=None, **kw):
|
||||
captured["url"] = url
|
||||
captured["payload"] = json
|
||||
return FakeResponse(json_data={"results": [
|
||||
{"title": "T", "url": "https://a.dev", "content": "c" * 800},
|
||||
]})
|
||||
|
||||
monkeypatch.setattr(web.httpx, "post", fake_post)
|
||||
results, engines = web._search_tavily("q", web.WebSearchInput(query="q", max_results=3))
|
||||
assert results[0]["title"] == "T"
|
||||
assert len(results[0]["snippet"]) <= 600
|
||||
assert engines == []
|
||||
assert captured["payload"]["api_key"] == ""
|
||||
assert captured["payload"]["max_results"] == 3
|
||||
|
||||
def test_brave_maps_results(self, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def fake_get(url, params=None, headers=None, **kw):
|
||||
captured["url"] = url
|
||||
captured["headers"] = headers
|
||||
return FakeResponse(json_data={"web": {"results": [
|
||||
{"title": "B", "url": "https://b.dev", "description": "d"},
|
||||
]}})
|
||||
|
||||
monkeypatch.setattr(web.httpx, "get", fake_get)
|
||||
results, _ = web._search_brave("q", web.WebSearchInput(query="q"))
|
||||
assert results[0]["title"] == "B"
|
||||
assert "api.search.brave.com" in str(captured["url"])
|
||||
|
||||
def test_serpapi_maps_results(self, monkeypatch):
|
||||
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
|
||||
"organic_results": [{"title": "S", "link": "https://s.dev", "snippet": "sn"}],
|
||||
}))
|
||||
results, _ = web._search_serpapi("q", web.WebSearchInput(query="q"))
|
||||
assert results[0]["url"] == "https://s.dev"
|
||||
|
||||
def test_exa_maps_results(self, monkeypatch):
|
||||
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: FakeResponse(json_data={
|
||||
"results": [{"title": "E", "url": "https://e.dev", "text": "t" * 900}],
|
||||
}))
|
||||
results, _ = web._search_exa("q", web.WebSearchInput(query="q"))
|
||||
assert results[0]["title"] == "E"
|
||||
assert len(results[0]["snippet"]) <= 600
|
||||
|
||||
|
||||
class TestProviderChain:
|
||||
def test_no_key_falls_back_to_searxng(self, monkeypatch):
|
||||
for var in ("OBSIGATE_TAVILY_API_KEY", "OBSIGATE_BRAVE_API_KEY",
|
||||
"OBSIGATE_SERPAPI_API_KEY", "OBSIGATE_EXA_API_KEY"):
|
||||
monkeypatch.delenv(var, raising=False)
|
||||
chain = [name for name, _ in web._provider_chain()]
|
||||
assert chain[0] == "searxng"
|
||||
assert "tavily" not in chain
|
||||
|
||||
def test_keyed_provider_used_first_when_key_set(self, monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
|
||||
chain = [name for name, _ in web._provider_chain()]
|
||||
assert chain[0] == "tavily"
|
||||
assert "brave" not in chain # no key → skipped
|
||||
|
||||
def test_explicit_order_env(self, monkeypatch):
|
||||
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
|
||||
monkeypatch.setenv("OBSIGATE_EXA_API_KEY", "k")
|
||||
monkeypatch.setenv("OBSIGATE_WEB_PROVIDERS", "exa,unknown,tavily")
|
||||
chain = [name for name, _ in web._provider_chain()]
|
||||
assert chain[:2] == ["exa", "tavily"]
|
||||
|
||||
def test_search_uses_keyed_provider_first(self, monkeypatch):
|
||||
_no_fallback(monkeypatch)
|
||||
monkeypatch.setenv("OBSIGATE_BRAVE_API_KEY", "k")
|
||||
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
|
||||
"web": {"results": [{"title": "B", "url": "https://b.dev", "description": "d"}]},
|
||||
}))
|
||||
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
|
||||
assert out["provider"] == "brave"
|
||||
assert out["count"] == 1
|
||||
|
||||
|
||||
class TestRetry:
|
||||
def test_transient_error_retried_then_succeeds(self, monkeypatch):
|
||||
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
|
||||
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
|
||||
calls = {"n": 0}
|
||||
|
||||
def flaky_get(*a, **kw):
|
||||
calls["n"] += 1
|
||||
if calls["n"] == 1:
|
||||
raise httpx.ConnectError("blip")
|
||||
return FakeResponse(json_data={"results": [
|
||||
{"title": "A", "url": "https://a.dev", "content": "x"}]})
|
||||
|
||||
monkeypatch.setattr(web.httpx, "get", flaky_get)
|
||||
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
|
||||
assert out["provider"] == "searxng"
|
||||
assert calls["n"] == 2
|
||||
|
||||
def test_persistent_error_not_retried_forever(self, monkeypatch):
|
||||
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
|
||||
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
|
||||
calls = {"n": 0}
|
||||
|
||||
def dead_get(*a, **kw):
|
||||
calls["n"] += 1
|
||||
raise httpx.ConnectError("down")
|
||||
|
||||
monkeypatch.setattr(web.httpx, "get", dead_get)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
web.web_search(_ctx(), web.WebSearchInput(query="q"))
|
||||
assert ei.value.code == "web_search_unavailable"
|
||||
assert calls["n"] == 2 # initial + 1 retry, per provider
|
||||
@@ -0,0 +1,73 @@
|
||||
"""Unit tests for the dynamic rendering path (#92): fetch_url(render=True)."""
|
||||
|
||||
import pytest
|
||||
|
||||
import backend.tools.web as web
|
||||
from backend.tools import webrender
|
||||
from backend.tools.context import ToolContext, ToolError, ToolMode
|
||||
from backend.tools.registry import get_tool
|
||||
|
||||
|
||||
def _ctx() -> ToolContext:
|
||||
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
|
||||
|
||||
|
||||
class TestRegistration:
|
||||
def test_render_param_exposed_in_schema(self):
|
||||
spec = get_tool("fetch_url")
|
||||
assert spec is not None
|
||||
assert "render" in spec.input_model.model_fields
|
||||
|
||||
|
||||
class TestRenderUnavailable:
|
||||
def test_missing_playwright_clear_error(self, monkeypatch):
|
||||
monkeypatch.setattr(webrender, "_playwright_available", lambda: False)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
web.fetch_url(_ctx(), web.FetchUrlInput(
|
||||
url="https://example.com/spa", render=True))
|
||||
assert ei.value.code == "playwright_unavailable"
|
||||
|
||||
def test_ssrf_guard_applied_before_render(self, monkeypatch):
|
||||
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
web.fetch_url(_ctx(), web.FetchUrlInput(
|
||||
url="http://127.0.0.1:9222/devtools", render=True))
|
||||
assert ei.value.code in ("ssrf_blocked", "dns_error")
|
||||
|
||||
|
||||
class TestRenderSuccess:
|
||||
def test_fetch_url_delegates_to_worker(self, monkeypatch):
|
||||
captured = {}
|
||||
|
||||
def fake_render(url):
|
||||
captured["url"] = url
|
||||
return {"url": url, "status": 200, "title": "SPA",
|
||||
"text": "dynamic content", "rendered": True, "truncated": False}
|
||||
|
||||
monkeypatch.setattr(webrender, "render_page", fake_render)
|
||||
out = web.fetch_url(_ctx(), web.FetchUrlInput(
|
||||
url="https://example.com/spa", render=True))
|
||||
assert captured["url"] == "https://example.com/spa"
|
||||
assert out["rendered"] is True
|
||||
assert "dynamic content" in out["text"]
|
||||
|
||||
def test_worker_failure_maps_to_tool_error(self, monkeypatch):
|
||||
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
|
||||
|
||||
def boom(url):
|
||||
raise RuntimeError("chromium crashed")
|
||||
|
||||
# The executor re-raises the worker exception on .result(); render_page
|
||||
# must wrap it into a ToolError instead of leaking a bare exception.
|
||||
monkeypatch.setattr(webrender, "_render_in_worker", boom)
|
||||
with pytest.raises(ToolError) as ei:
|
||||
web.fetch_url(_ctx(), web.FetchUrlInput(
|
||||
url="https://example.com/spa", render=True))
|
||||
assert ei.value.code == "render_unavailable"
|
||||
|
||||
|
||||
class TestMarkdownExtraction:
|
||||
def test_html_to_text_reused(self):
|
||||
text = webrender._html_to_text("<html><body><p>hello</p><script>x()</script></body></html>")
|
||||
assert "hello" in text
|
||||
assert "x()" not in text
|
||||
Reference in New Issue
Block a user