Compare commits

...
10 Commits
Author SHA1 Message Date
bruno 2e2a33cef3 fix: corrige 6 bugs mineurs (BUG-035 a BUG-040)
CI / lint (push) Successful in 1m36s
CI / security (push) Successful in 1m4s
CI / test (push) Successful in 3m41s
CI / build (push) Successful in 59s
CI / e2e (push) Successful in 11m8s
2026-09-17 20:05:08 -04:00
bruno 2c460022f8 test(pdf): helper d'ouverture robuste au vault replie (BUG-060)
CI / lint (push) Successful in 1m37s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 3m33s
CI / build (push) Successful in 59s
CI / e2e (push) Successful in 10m57s
2026-09-17 19:42:43 -04:00
bruno 133644a0ba fix(pdf): affichage des pages du viewer PDF via iframe (BUG-060)
CI / lint (push) Successful in 1m36s
CI / security (push) Successful in 1m3s
CI / test (push) Failing after 3m41s
CI / build (push) Skipped
CI / e2e (push) Skipped
2026-09-17 19:39:48 -04:00
bruno ba0ec3d1fa feat(config): cles des sources connectees et recherche a cle editables depuis la page Configurations (#103)
CI / lint (push) Successful in 1m46s
CI / security (push) Successful in 1m23s
CI / test (push) Successful in 3m42s
CI / build (push) Successful in 58s
CI / e2e (push) Successful in 10m50s
2026-09-17 13:59:26 -04:00
bruno 22e9240e4f fix(ai): les clics ne font plus sauter la conversation au bas (BUG-059); tableaux markdown corrects dans create_pdf (#92) 2026-09-17 13:55:36 -04:00
bruno 6a58a59a11 feat(ai): ecosysteme d'outils phase 2 - recherche a cle, cache/retry, Playwright, crawl, Gitea/GitHub, documents (#92)
CI / lint (push) Successful in 1m37s
CI / security (push) Successful in 1m1s
CI / test (push) Successful in 3m26s
CI / build (push) Successful in 1m44s
CI / e2e (push) Successful in 11m1s
2026-09-17 11:52:03 -04:00
bruno 26328fadeb fix(editor): la barre de numerotation de ligne suit le theme (BUG-058)
CI / lint (push) Successful in 1m37s
CI / security (push) Successful in 1m1s
CI / test (push) Successful in 2m54s
CI / build (push) Successful in 56s
CI / e2e (push) Successful in 10m43s
2026-09-17 10:02:48 -04:00
bruno c0eea526de feat(ai): ajouter une section de la reponse et supporter l'editeur Forge (BUG-057, #102)
CI / lint (push) Successful in 1m34s
CI / security (push) Successful in 1m1s
CI / test (push) Successful in 2m44s
CI / build (push) Successful in 1m1s
CI / e2e (push) Successful in 10m51s
L'assistant IA peut desormais inserer sa reponse dans l'editeur Forge (postMessage parent-insert) en plus d'Editer, et chaque bloc de code propose un bouton « Ajouter la section » pour n'inserer que ce bloc.
2026-09-17 09:39:35 -04:00
bruno 94ea5909f4 fix(forge): le parent quitte aussi le plein ecran avant l'assistant IA (BUG-056)
CI / lint (push) Successful in 1m34s
CI / security (push) Successful in 1m1s
CI / test (push) Successful in 2m11s
CI / build (push) Successful in 57s
CI / e2e (push) Successful in 10m50s
2026-09-17 09:19:06 -04:00
bruno 3758db2861 fix(forge): quitter le plein ecran avant d'ouvrir l'assistant IA (BUG-056)
CI / lint (push) Successful in 1m33s
CI / security (push) Successful in 1m0s
CI / test (push) Successful in 3m7s
CI / build (push) Successful in 57s
CI / e2e (push) Successful in 10m56s
2026-09-17 09:07:00 -04:00
63 changed files with 4104 additions and 114 deletions
+21
View File
@@ -7,6 +7,11 @@ OBSIGATE_AUTH_ENABLED=true
OBSIGATE_ADMIN_USER=admin
OBSIGATE_ADMIN_PASSWORD=chab30
# DANGER : si OBSIGATE_AUTH_ENABLED=false, toute requête devient un admin
# anonyme. Le serveur REFUSE de démarrer sur une adresse non-loopback
# (ex. 0.0.0.0) sauf si l'on force l'opt-in ci-dessous. À réserver au local.
# OBSIGATE_ALLOW_INSECURE=false
# Sécurité des cookies (activer si derrière HTTPS)
# OBSIGATE_SECURE_COOKIES=false
@@ -74,3 +79,19 @@ DEEPSEEK_MODEL=deepseek-chat
# Chaîne de repli sans clé (DuckDuckGo puis Bing) si SearXNG ne remonte rien
# OBSIGATE_WEB_FALLBACK=1
# OBSIGATE_WEB_TIMEOUT=10
# Fournisseurs à clé (#92), essayés avant SearXNG — injecter via Infisical en prod
# OBSIGATE_TAVILY_API_KEY=
# OBSIGATE_BRAVE_API_KEY=
# OBSIGATE_SERPAPI_API_KEY=
# OBSIGATE_EXA_API_KEY=
# Ordre des fournisseurs (sinon : clés présentes puis SearXNG puis replis)
# OBSIGATE_WEB_PROVIDERS=brave,searxng
# Réessais réseau (backoff maison) + cache SQLite des résultats web
# OBSIGATE_WEB_RETRY=1
# OBSIGATE_WEB_CACHE_TTL=900 # secondes ; 0 = cache désactivé
# Rendu dynamique (pages SPA) — dépendance optionnelle :
# pip install playwright && playwright install chromium
# ── Assistant IA — sources connectées (Gitea / GitHub) ──
# OBSIGATE_GITEA_URL=https://git.example.net
# OBSIGATE_GITEA_TOKEN=
# OBSIGATE_GITHUB_TOKEN=
+2
View File
@@ -38,6 +38,7 @@ jobs:
- name: Frontend unit tests
run: |
node tests/frontend/unit.test.mjs
node tests/frontend/pdf-viewer.test.mjs
node tests/frontend/forge-completion.test.mjs
- name: Frontend JSDOM tests (PaneManager + Excalidraw + Plugins + AI + SW + Collab + Mobile + Semantic + Desktop + Inline edition)
@@ -196,6 +197,7 @@ jobs:
-e DIR_1_NAME=TestDir \
-e DIR_1_PATH=/vaults/TestDir \
-e OBSIGATE_AUTH_ENABLED=false \
-e OBSIGATE_ALLOW_INSECURE=true \
obsigate:ci
# Docker-in-docker : le bind mount $(pwd)/... pointe sur un chemin
# du job container, inexistant sur l'hôte → montage vide. Les -v
+193 -1
View File
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
> **En cours de développement** : les changements à venir sont listés dans la section
> [Unreleased](#unreleased). La dernière version livrée est **2.8.2**.
> [Unreleased](#unreleased). La dernière version livrée est **2.11.3**.
---
@@ -14,6 +14,198 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
---
## [2.11.3] — 2026-09-17
### Corrigé
- **BUG-035 — `secret_redactor` : faux positifs sur les hashs hex** : la règle qui masquait
tout jeton hexadécimal de 40 à 64 caractères mutilait les hashs git/SHA légitimes des notes.
Le masquage des chaînes hexadécimales n'a désormais lieu que si un mot-clé de secret
(`secret`, `token`, `key`, `password`, `bearer`…) figure dans les 60 caractères précédents ;
un contexte de hash (`commit`, `sha256`, `hash`, `checksum`, `git`, `etag`…) exempte
explicitement la chaîne. Fichier : `backend/secret_redactor.py`.
- **BUG-036 — Collaboration WebSocket : jeton accepté en query string** : le JWT n'est plus lu
depuis `?token=` (URLs journalisées par les proxies et l'historique navigateur). Le cookie
HttpOnly `access_token`, envoyé automatiquement par le navigateur lors du handshake
same-origin, est le seul transport supporté ; les trames brutes dépassant
`MAX_MESSAGE_CHARS` (16 Mio) sont rejetées avant analyse. Fichier : `backend/collab.py`.
- **BUG-037 — Mode sans authentification** : au démarrage, un avertissement explicite est
journalisé quand `OBSIGATE_AUTH_ENABLED=false`. Le serveur **refuse désormais de démarrer**
s'il est lié à une adresse non-loopback sans l'opt-in explicite `OBSIGATE_ALLOW_INSECURE=true`,
pour empêcher l'exposition publique d'une instance sans authentification (admin anonyme).
Fichiers : `backend/auth/middleware.py`, `backend/main.py`.
- **BUG-038 — Argon2 : coût mémoire recalibré** : `memory_cost` passe de 64 Mio à 19 Mio
(`m=19456 Kio, t=2, p=1`, recommandation OWASP actuelle) pour supprimer le risque
d'épuisement mémoire sous connexions simultanées ; les anciens hachages restent valides et
sont migrés automatiquement (`needs_rehash`). Fichier : `backend/auth/password.py`.
- **BUG-039 — Énumération de comptes au login** : les comptes inconnus, désactivés, verrouillés
et limités par le budget par compte répondent tous un `401 Identifiants invalides` avec un
temps équivalent (hachage factice), au lieu d'un `429`/`403` distinctif ; seul le rate-limit
par IP (non lié à un compte) conserve le `429`. Fichier : `backend/auth/router.py`.
- **BUG-040 — Extraction PDF différée au scan** : `_scan_vault` ne lit plus que les métadonnées
des PDF ; l'extraction de texte intégrale (100 kio) est déléguée à `enrich_pdf_texts()`,
exécutée après la construction de l'index/inverted index (démarrage) et après chaque
réindexation. Un vault contenant de nombreux/gros PDF démarre sans être bloqué ; le texte
reste recherchable une fois l'enrichissement terminé. Fichiers : `backend/indexer.py`,
`backend/main.py`.
- **Tests** : `tests/test_api_main.py` (redactor hex), `tests/test_auth.py` (coût Argon2,
garde-fou d'instance non authentifiée), `tests/test_auth_api.py` (login uniforme),
`tests/test_collab.py` (jeton query rejeté, trame surdimensionnée), `tests/test_pdf.py`
(scan différé + enrichissement).
---
## [2.11.2] — 2026-09-17
### Modifié
- **Tests E2E PDF (BUG-060)** : le helper d'ouverture de fichier de
`tests/e2e/pdf-viewer.spec.js` étend désormais le vault dans l'arborescence s'il est replié
(cas d'une instance avec authentification activée) au lieu de supposer le vault déjà déplié.
---
## [2.11.1] — 2026-09-17
### Corrigé
- **BUG-060 — Viewer PDF : l'affichage des pages ne fonctionnait pas** : au clic sur un fichier
`.pdf`, seule la barre d'outils « PDF — N pages » s'affichait, le corps restant vide. La CSP
durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>`, alors
que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était
bloqué. Le rendu passe désormais par une `<iframe>` (autorisée par `frame-src 'self'`, le stream
`/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` est conservé. Fichiers :
`frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau),
`tests/e2e/pdf-viewer.spec.js` (nouveau, fixture `test_vault/sample-pdf.pdf`).
---
## [2.11.0] — 2026-09-17
### Ajouté
- **#103 — Clés des sources connectées configurables depuis la page Configurations** : nouvelle
section « 🔗 Sources connectées & recherche » (admin) permettant de saisir ses propres jetons
sans toucher aux variables d'environnement : clés de recherche web à clé (Tavily, Brave,
SerpAPI, Exa), URL Gitea + token, token GitHub. Stockage dans `data/api_keys.json` (même
fichier que les clés des fournisseurs IA) via les endpoints
`GET/POST/DELETE /api/config/tool-keys` (GET masque les secrets, URLs en clair ; liste
blanche stricte ; admin requis). Les outils lisent la valeur stockée **en priorité** puis
l'environnement (Infisical) — `backend/tools/secrets.py`. Fichiers :
`backend/tools/{secrets,web,connected}.py`, `backend/main.py`, `frontend/index.html`,
`frontend/js/config.js`, `frontend/locales/{fr,en}.json`, `tests/test_tool_keys.py` (nouveau),
`tests/test_connected_tools.py`.
---
## [2.10.1] — 2026-09-17
### Corrigé
- **BUG-059 — Assistant IA : un clic dans la conversation faisait sauter le texte au bas de la
fenêtre** : dans une conversation ouverte (post ancré en haut), tout clic — lien de fichier,
étapes, sélection de texte — dépinait l'ancre via le gestionnaire `mousedown` et retirait le
padding d'ancre, bornant le scroll à la nouvelle hauteur max (saut au bas). Le dépintage est
désormais réservé aux vrais gestes de scroll : molette, tactile et poignée de scroll
uniquement (`isScrollbarPress`). Fichiers : `frontend/js/bookslm.js`,
`tests/frontend/ai.test.mjs`.
- **#92 — `create_pdf` de l'Assistant : tableaux mal formatés** : l'outil `create_pdf`
(génération PDF via l'assistant) utilisait un rendu reportlab simplifié sans support des
tableaux. Il passe désormais par le même pipeline que le bouton « Télécharger PDF » de la
page document (mistune + plugin `table` + WeasyPrint), avec repli automatique sur le rendu
simple si WeasyPrint n'est pas disponible (GTK absent). Fichiers :
`backend/tools/documents.py`, `tests/test_document_tools.py`.
---
## [2.10.0] — 2026-09-17
### Ajouté
- **#92 — Assistant IA : écosystème d'outils phase 2 (web étendu, sources connectées, documents)** :
- **Recherche web à clé** : fournisseurs optionnels essayés avant SearXNG — Tavily, Brave
Search, SerpAPI (Google) et Exa (`OBSIGATE_TAVILY_API_KEY`, `OBSIGATE_BRAVE_API_KEY`,
`OBSIGATE_SERPAPI_API_KEY`, `OBSIGATE_EXA_API_KEY`), avec ordre personnalisable via
`OBSIGATE_WEB_PROVIDERS`.
- **Transverse** : réessais réseau avec backoff maison (`OBSIGATE_WEB_RETRY`) et cache SQLite
des résultats web avec TTL (`OBSIGATE_WEB_CACHE_TTL`, 900 s par défaut, `OBSIGATE_WEB_CACHE_PATH`).
- **Rendu dynamique** : `fetch_url(render=True)` délègue les pages SPA à un worker Playwright
isolé (dépendance optionnelle, dégradation propre si non installée).
- **Crawl de site** : `crawl_site` (WRITE, confirmation) capture jusqu'à 20 pages d'un même
hôte et enregistre un condensé Markdown dans un vault.
- **Sources connectées** : Gitea (`OBSIGATE_GITEA_URL`/`OBSIGATE_GITEA_TOKEN`) et GitHub
(`OBSIGATE_GITHUB_TOKEN`) — `git_list_repos`, `git_search_issues`, `git_get_file` (READ,
rate-limités, audités). Les drives cloud (Google Drive / OneDrive) restent orientés serveur
MCP externe (#79), conformément à la feuille de route.
- **Production de documents** : `create_xlsx` (openpyxl), `create_docx` (python-docx),
`create_csv`, `create_pdf` (reportlab) — outils WRITE avec confirmation et sauvegarde
dans le vault (backup avant écrasement).
- Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests unitaires mockés
(httpx). Dépendances ajoutées : `openpyxl`, `python-docx`, `reportlab`.
Fichiers : `backend/tools/{webcache,webrender,connected,crawler,documents,web,schemas,labels}.py`,
`backend/services/mutations.py`, `tests/test_{web_cache,web_search_providers,webrender,connected_tools,document_tools,crawler}.py`.
---
## [2.9.1] — 2026-09-17
### Corrigé
- **BUG-058 — Éditeur « Editer » : la barre de numérotation de ligne ne suit pas le thème** :
CodeMirror peint son gutter (colonne des numéros de ligne) avec des valeurs claires codées
en dur (`#f5f5f5`, bordure `#ddd`), si bien qu'en thème sombre la barre restait gris clair
alors que le fond de l'éditeur suivait le thème. Le gutter est désormais dérivé des
variables CSS du thème ObsiGate (`color-mix(in srgb, var(--text-primary) …)` pour un fond
subtil, `--text-secondary` pour les numéros, `--border` pour la séparation), ce qui le fait
suivre tous les thèmes et modes (sombre, clair, contraste élevé, sépia). Fichiers :
`frontend/style.css`. Tests : `tests/frontend/editor-inline.test.mjs`.
---
## [2.9.0] — 2026-09-17
### Ajouté
- **#102 — Assistant IA : bouton « Ajouter la section » par bloc de code** : chaque bloc
de code d'une réponse de l'assistant (par ex. une section ```markdown```) affiche un
bouton discret « Ajouter la section » qui insère **uniquement ce bloc** dans le document
ouvert (sans les délimiteurs de code), au lieu de la réponse complète. Fichiers :
`frontend/js/bookslm.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`.
Tests : `tests/frontend/ai.test.mjs`.
### Corrigé
- **BUG-057 — Assistant IA : bouton « Ajouter » inopérant dans l'éditeur Forge** : le
bouton ne ciblait que `state.editorView` (CodeMirror de « Editer ») et affichait « Aucun
document ouvert dans l'éditeur » en Forge. `_insertIntoEditor()` prend désormais en
charge les trois surfaces : CodeMirror, l'iframe Forge (délégation par
`postMessage({ type: 'parent-insert' })` → `insertAtCursor` dans `editor-poc.html`) et le
textarea de repli. Fichiers : `frontend/js/bookslm.js`, `frontend/editor-poc.html`.
Tests : `tests/frontend/ai.test.mjs`, `tests/frontend/editor-inline.test.mjs`.
---
## [2.8.4] — 2026-09-17
---
## [2.8.3] — 2026-09-17
### Corrigé
- **BUG-056 - Éditeur Forge en plein écran : l'Assistant IA s'ouvre en arrière-plan** :
le panneau de l'assistant est monté dans le document parent, alors que le plein écran
Forge porte sur l'iframe — l'API Fullscreen n'affichant que l'élément plein écran et ses
descendants, le panneau restait invisible. Le plein écran est désormais quitté avant
d'ouvrir le panneau : côté **parent** (`sync.js`, sur `forge-open-ai` — le plein écran
peut appartenir au document parent et non à l'iframe) **et** côté iframe
(`editor-poc.html`, `openAssistant`), la demande d'ouverture étant émise une fois la
sortie effective. Fichiers : `frontend/editor-poc.html`, `frontend/js/sync.js`.
Tests : `tests/frontend/forge-completion.test.mjs` (+1),
`tests/frontend/editor-inline.test.mjs` (+1).
---
## [2.8.2] — 2026-09-17
### Corrigé
+13 -3
View File
@@ -4,7 +4,7 @@
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
[![Version](https://img.shields.io/badge/Version-2.8.2-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.11.3-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -282,6 +282,16 @@ Un compte **admin** connecté voit une icône 🛡️ dans le header : liste, cr
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Autoriser les webhooks vers des adresses privées/boucle | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Taille max des PDF extraits (text indexation) | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | Timeout extraction PDF (secondes) | `30` |
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Fournisseurs de recherche web à clé (essayés avant SearXNG) | — |
| `OBSIGATE_WEB_PROVIDERS` | Ordre des fournisseurs de recherche (ex. `brave,searxng`) | — |
| `OBSIGATE_WEB_RETRY` | Réessais réseau des outils web (backoff maison) | `1` |
| `OBSIGATE_WEB_CACHE_TTL` | Durée du cache SQLite des résultats web (secondes, `0` = off) | `900` |
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Source connectée Gitea (outil `git_list_repos`…) | — |
| `OBSIGATE_GITHUB_TOKEN` | Jeton GitHub (outil `git_list_repos`…) | — |
> Ces clés peuvent aussi être saisies **depuis l'interface** (menu → Configurations →
> « Sources connectées & recherche ») : la valeur saisie est stockée dans `data/api_keys.json`
> et prime sur la variable d'environnement.
### Volume pour la persistance
@@ -916,8 +926,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
## 📝 Changelog
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.8.2).
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.11.3).
---
*Projet : ObsiGate | Version : 2.8.2 | Dernière mise à jour : Juin 2026*
*Projet : ObsiGate | Version : 2.11.3 | Dernière mise à jour : Juin 2026*
+13 -3
View File
@@ -2,7 +2,7 @@
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
[![Version](https://img.shields.io/badge/Version-2.8.2-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.11.3-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -320,6 +320,16 @@ When an **admin** account is logged in, a 🛡️ icon appears in the header. Cl
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Allow webhooks to private/loopback addresses | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Max PDF size for text extraction | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | PDF extraction timeout (seconds) | `30` |
| `OBSIGATE_TAVILY_API_KEY` / `OBSIGATE_BRAVE_API_KEY` / `OBSIGATE_SERPAPI_API_KEY` / `OBSIGATE_EXA_API_KEY` | Keyed web-search providers (tried before SearXNG) | — |
| `OBSIGATE_WEB_PROVIDERS` | Search provider order (e.g. `brave,searxng`) | — |
| `OBSIGATE_WEB_RETRY` | Web tools network retries (house-made backoff) | `1` |
| `OBSIGATE_WEB_CACHE_TTL` | SQLite cache TTL for web results (seconds, `0` = off) | `900` |
| `OBSIGATE_GITEA_URL` / `OBSIGATE_GITEA_TOKEN` | Gitea connected source (`git_list_repos`…) | — |
| `OBSIGATE_GITHUB_TOKEN` | GitHub token (`git_list_repos`…) | — |
> These keys can also be entered **from the UI** (menu → Configurations →
> "Connected sources & search"): the stored value goes to `data/api_keys.json`
> and takes precedence over the environment variable.
>All these variables are documented in `.env.example`.
@@ -1085,8 +1095,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
## 📝 Changelog
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.8.2).
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.11.3).
---
*Project: ObsiGate | Version: 2.8.2 | Last updated: May 2026*
*Project: ObsiGate | Version: 2.11.3 | Last updated: May 2026*
+1 -1
View File
@@ -1 +1 @@
2.8.2
2.11.3
+32
View File
@@ -4,6 +4,7 @@
import logging
import os
import sys
from fastapi import Depends, HTTPException, Request
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
@@ -17,6 +18,9 @@ logger = logging.getLogger("obsigate.auth.middleware")
security = HTTPBearer(auto_error=False)
#: Hosts considered safe to bind without authentication (loopback only).
_LOOPBACK_HOSTS = {"127.0.0.1", "::1", "localhost", "0:0:0:0:0:0:0:1"}
def is_auth_enabled() -> bool:
"""Check if authentication is enabled via environment variable.
@@ -26,6 +30,34 @@ def is_auth_enabled() -> bool:
return os.environ.get("OBSIGATE_AUTH_ENABLED", "true").lower() != "false"
def is_insecure_mode_allowed() -> bool:
"""True when the operator explicitly accepts running without auth (BUG-037)."""
return os.environ.get("OBSIGATE_ALLOW_INSECURE", "false").lower() in ("1", "true", "yes", "on")
def bind_host_from_argv(argv: list[str] | None = None) -> str | None:
"""Extract the ``--host`` value from the process arguments (uvicorn), if any.
Returns ``None`` when no explicit host is passed (uvicorn then defaults to
loopback ``127.0.0.1``).
"""
args = sys.argv if argv is None else argv
for i, arg in enumerate(args):
if arg == "--host" and i + 1 < len(args):
return args[i + 1]
if arg.startswith("--host="):
return arg.split("=", 1)[1]
return None
def is_loopback_host(host: str | None) -> bool:
"""True when *host* is a loopback address (or unset → uvicorn default)."""
if not host:
return True
normalized = host.strip().strip("[]").lower()
return normalized in _LOOPBACK_HOSTS
def get_current_user(
request: Request,
credentials: HTTPAuthorizationCredentials | None = Depends(security),
+11 -4
View File
@@ -1,14 +1,21 @@
# backend/auth/password.py
# Argon2id password hashing — OWASP 2024 recommended algorithm.
# Parameters: time_cost=2, memory_cost=64MB, parallelism=2
# Parameters (BUG-038): time_cost=2, memory_cost=19 MiB, parallelism=1
# (OWASP current recommendation for Argon2id). The previous 64 MiB setting
# allowed memory exhaustion under concurrent login attempts.
from argon2 import PasswordHasher
from argon2.exceptions import VerificationError, VerifyMismatchError
#: Argon2id cost parameters (OWASP 2024: m=19456 KiB, t=2, p=1).
ARGON2_TIME_COST = 2
ARGON2_MEMORY_COST_KIB = 19456 # 19 MiB
ARGON2_PARALLELISM = 1
ph = PasswordHasher(
time_cost=2,
memory_cost=65536, # 64 MB
parallelism=2,
time_cost=ARGON2_TIME_COST,
memory_cost=ARGON2_MEMORY_COST_KIB,
parallelism=ARGON2_PARALLELISM,
hash_len=32,
salt_len=16,
)
+18 -17
View File
@@ -124,31 +124,32 @@ async def auth_status():
async def login(body: LoginRequest, response: Response, request: Request):
"""Authenticate a user. Returns access token and sets refresh cookie.
Implements timing-safe responses to prevent user enumeration:
a failed login with an unknown user takes the same time as one
with a known user (dummy hash is computed).
Implements timing-safe responses to prevent user enumeration: a failed
login with an unknown user takes the same time as one with a known user
(dummy hash is computed). BUG-039: unknown, inactive, locked and
per-account rate-limited accounts all answer the same ``401`` so the HTTP
status can never reveal whether an account exists.
"""
client_ip = get_client_ip(request)
# IP-based rate limiting (10 failures / 15 min per IP). It is not
# account-specific, so a 429 here cannot be used to enumerate accounts.
if is_rate_limited(client_ip):
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
user = get_user(body.username)
if not user:
# BUG-039: uniform 401 + equivalent timing for every account-state outcome.
if not user or not user.get("active"):
# Timing-safe: simulate hash computation to prevent user enumeration
hash_password("dummy_timing_protection")
raise HTTPException(401, "Identifiants invalides")
if not user.get("active"):
raise HTTPException(403, "Compte désactivé")
# IP-based rate limiting (10 failures / 15 min per IP)
client_ip = get_client_ip(request)
if is_rate_limited(client_ip):
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
# BUG-031: per-account budget still applies when the attacker rotates IPs.
if is_account_rate_limited(body.username):
raise HTTPException(429, "Trop de tentatives sur ce compte (15min)")
if is_locked(body.username):
raise HTTPException(429, "Compte temporairement verrouillé (15min)")
# Kept indistinguishable from a wrong password (BUG-039).
if is_account_rate_limited(body.username) or is_locked(body.username):
hash_password("dummy_timing_protection")
raise HTTPException(401, "Identifiants invalides")
if not verify_password(body.password, user["password_hash"]):
attempts = record_login_failure(body.username)
+14 -4
View File
@@ -44,6 +44,9 @@ MAX_UPDATE_BYTES = 8 * 1024 * 1024
#: Taille maximale d'un snapshot texte (protection anti-abus).
MAX_TEXT_CHARS = 8 * 1024 * 1024
#: Taille maximale d'un message brut reçu (protection anti-abus, BUG-036).
MAX_MESSAGE_CHARS = 16 * 1024 * 1024
#: Palette de couleurs attribuées aux utilisateurs (curseurs + avatars).
PEER_COLORS = [
"#e6194b", "#3cb44b", "#4363d8", "#f58231", "#911eb4",
@@ -68,9 +71,13 @@ def authenticate_websocket(websocket: WebSocket) -> dict[str, Any] | None:
"""Authenticate a WebSocket connection.
Mirrors :func:`backend.auth.middleware.get_current_user` but works on the
WebSocket scope: the JWT is read from the ``access_token`` cookie (sent
automatically by same-origin browsers during the handshake) or, as a
fallback, from the ``token`` query parameter.
WebSocket scope: the JWT is read from the ``access_token`` cookie, which
same-origin browsers send automatically during the handshake.
BUG-036: the token is **never** accepted from the query string anymore —
URLs end up in access logs, proxies and browser history. Browsers cannot
set custom headers on a WebSocket handshake, so the HttpOnly cookie set at
login is the only supported transport.
Returns the user dict, or ``None`` if authentication fails.
"""
@@ -88,7 +95,7 @@ def authenticate_websocket(websocket: WebSocket) -> dict[str, Any] | None:
"_token_vaults": ["*"],
}
token = websocket.query_params.get("token") or websocket.cookies.get("access_token")
token = websocket.cookies.get("access_token")
if not token:
return None
@@ -274,6 +281,9 @@ class CollabManager:
# -- message handling ---------------------------------------------------
async def _on_message(self, room: CollabRoom, client: CollabClient, raw: str) -> None:
# BUG-036: drop oversized frames before parsing them.
if not isinstance(raw, str) or len(raw) > MAX_MESSAGE_CHARS:
return
try:
message = json.loads(raw)
except (ValueError, TypeError):
+74 -6
View File
@@ -481,12 +481,18 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
# PDF handling — special path (binary, uses pdf_reader)
tags: list[str] = []
pdf_text_pending = False
if ext == ".pdf":
from backend.pdf_reader import extract_pdf_metadata, extract_pdf_text
raw = extract_pdf_text(fpath, max_chars=100000)
from backend.pdf_reader import extract_pdf_metadata
# BUG-040: only the (cheap) metadata is read during the
# scan. Full-text extraction is deferred to a background
# pass (``enrich_pdf_texts``) so a vault with many/large
# PDFs no longer blocks startup and index rebuilds.
pdf_meta = extract_pdf_metadata(fpath)
title = pdf_meta.get("title") or fpath.stem.replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
raw = ""
content_preview = ""
pdf_text_pending = True
elif ext == ".excalidraw" or fpath.name.lower().endswith(".excalidraw.md"):
raw = fpath.read_text(encoding="utf-8", errors="replace")
raw = extract_excalidraw_indexable(raw)
@@ -510,7 +516,7 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
title, post.content
)
files.append({
file_info = {
"path": str(relative).replace("\\", "/"),
"title": title,
"tags": tags,
@@ -519,7 +525,10 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
"size": stat.st_size,
"modified": modified,
"extension": ext,
})
}
if pdf_text_pending:
file_info["pdf_text_pending"] = True
files.append(file_info)
for tag in tags:
tag_counts[tag] = tag_counts.get(tag, 0) + 1
@@ -535,6 +544,60 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
return {"files": files, "tags": tag_counts, "path": vault_path, "paths": paths, "config": {}}
async def enrich_pdf_texts(vault_name: str | None = None) -> int:
"""Extract text from PDFs whose extraction was deferred during the scan (BUG-040).
``_scan_vault`` only reads PDF metadata so a vault with many or large PDFs
starts serving immediately. This coroutine runs *after* the index (and the
inverted index) is ready, extracts the missing text off the event loop and
updates the in-memory entry plus the incremental index hooks.
Args:
vault_name: Restrict the pass to a single vault; ``None`` covers every
indexed vault.
Returns:
Number of deferred PDFs whose text extraction was attempted.
"""
from backend.pdf_reader import extract_pdf_text
pending: list[tuple[str, dict[str, Any], Path]] = []
with _index_lock:
for name, vault_data in index.items():
if vault_name is not None and name != vault_name:
continue
vault_root = Path(vault_data.get("path", ""))
for file_info in vault_data.get("files", []):
if file_info.get("pdf_text_pending"):
pending.append((name, file_info, vault_root / file_info["path"]))
if not pending:
return 0
loop = asyncio.get_running_loop()
enriched = 0
for name, file_info, file_path in pending:
try:
raw = await loop.run_in_executor(None, extract_pdf_text, file_path, 100000)
except Exception as exc: # pragma: no cover - defensive
logger.warning("PDF enrichment failed for %s: %s", file_path, exc)
raw = ""
file_info["content"] = raw[:SEARCH_CONTENT_LIMIT]
file_info["content_preview"] = raw[:200].strip()
file_info.pop("pdf_text_pending", None)
enriched += 1
if _on_index_change:
try:
_on_index_change("add", name, file_info["path"], file_info)
except Exception as exc: # pragma: no cover - defensive
logger.warning(
"Index hook failed after PDF enrichment for %s: %s", file_path, exc
)
logger.info("PDF enrichment: extracted text for %d deferred PDF(s)", enriched)
return enriched
async def build_index(progress_callback=None) -> None:
"""Build the full in-memory index for all configured vaults.
@@ -632,6 +695,8 @@ async def reload_index() -> dict[str, Any]:
Dict mapping vault names to their file/tag counts.
"""
await build_index()
# BUG-040: complete the deferred PDF extraction for the rebuilt index.
await enrich_pdf_texts()
stats = {}
for name, data in index.items():
stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])}
@@ -695,7 +760,10 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]:
# Rebuild attachment index for this vault only
from backend.attachment_indexer import build_attachment_index
await build_attachment_index({vault_name: config})
# BUG-040: complete the deferred PDF extraction for this vault.
await enrich_pdf_texts(vault_name)
stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])}
logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags")
return stats
+112 -1
View File
@@ -722,12 +722,56 @@ class SecurityHeadersMiddleware(BaseHTTPMiddleware):
return response
def _guard_insecure_auth() -> None:
"""Warn or refuse to start when authentication is disabled (BUG-037).
With ``OBSIGATE_AUTH_ENABLED=false`` every request is served as an
anonymous admin. That is convenient for local use but dangerous when the
process is reachable from a network. Binding to a non-loopback host
without the explicit ``OBSIGATE_ALLOW_INSECURE=true`` opt-in is refused.
"""
from backend.auth.middleware import (
bind_host_from_argv,
is_auth_enabled,
is_insecure_mode_allowed,
is_loopback_host,
)
if is_auth_enabled():
return
if is_insecure_mode_allowed():
logger.warning(
"Authentication is DISABLED and OBSIGATE_ALLOW_INSECURE=true: every request "
"is treated as an anonymous administrator. Do not expose this instance."
)
return
host = bind_host_from_argv()
if not is_loopback_host(host):
raise RuntimeError(
"Refusing to start: authentication is disabled (OBSIGATE_AUTH_ENABLED=false) "
f"while binding to a non-loopback address ('{host}'). This would expose an "
"unauthenticated instance with admin access. Enable authentication, or set "
"OBSIGATE_ALLOW_INSECURE=true if you really know what you are doing."
)
logger.warning(
"Authentication is DISABLED (OBSIGATE_AUTH_ENABLED=false): every request is "
"treated as an anonymous administrator. This is only safe on a trusted, "
"loopback-only deployment."
)
@asynccontextmanager
async def lifespan(app: FastAPI):
"""Application lifespan: build index on startup, cleanup on shutdown."""
global _search_executor, _vault_watcher
_search_executor = ThreadPoolExecutor(max_workers=2, thread_name_prefix="search")
# BUG-037: refuse to expose an unauthenticated instance on a public bind.
_guard_insecure_auth()
# Bootstrap admin account if needed
bootstrap_admin()
@@ -748,6 +792,11 @@ async def lifespan(app: FastAPI):
# Build the semantic (embedding) index in the same background thread pool.
await loop.run_in_executor(_search_executor, init_semantic_index)
# BUG-040: extract the PDF text deferred during the scan now that the
# index and inverted index are queryable (keeps startup non-blocking).
from backend.indexer import enrich_pdf_texts
await enrich_pdf_texts()
# Scan for plugins in all vaults
logger.info("Scanning for plugins...")
from backend.indexer import vault_config
@@ -3679,6 +3728,68 @@ async def api_delete_ai_key(provider_env: str, current_user=Depends(require_admi
return {"status": "deleted", "key": key_name}
# ---------------------------------------------------------------------------
# Tool & connected-source keys (#103) — same store as the AI provider keys
# ---------------------------------------------------------------------------
from backend.tools.secrets import (
TOOL_KEY_NAMES as _TOOL_KEY_NAMES,
)
from backend.tools.secrets import (
delete_tool_key as _delete_tool_key,
)
from backend.tools.secrets import (
get_tool_key as _get_tool_key,
)
from backend.tools.secrets import (
mask_value as _mask_tool_value,
)
from backend.tools.secrets import (
set_tool_key as _set_tool_key,
)
@app.get("/api/config/tool-keys", response_model=AIKeysResponse)
async def api_get_tool_keys(current_user=Depends(require_admin)):
"""Return tool/connected-source configuration (tokens masked, URLs clear)."""
masked = {}
for name in _TOOL_KEY_NAMES:
masked[name] = _mask_tool_value(name, _get_tool_key(name))
return masked
@app.post("/api/config/tool-keys", response_model=StatusResponse)
async def api_set_tool_keys(body: dict = Body(...), current_user=Depends(require_admin)):
"""Save tool/connected-source keys.
Only whitelisted names (``backend.tools.secrets.TOOL_KEY_NAMES``) are
accepted: Tavily/Brave/SerpAPI/Exa API keys, Gitea URL + token, GitHub
token. Empty values delete the stored entry.
"""
updated = []
for name, value in body.items():
if name not in _TOOL_KEY_NAMES:
raise HTTPException(status_code=400, detail=f"Clé inconnue: {name}")
if value is not None and not isinstance(value, str):
raise HTTPException(status_code=400, detail=f"Type invalide pour {name}")
_set_tool_key(name, value or "")
updated.append(name)
logger.info(f"Tool keys updated: {updated}")
return {"status": "ok"}
@app.delete("/api/config/tool-keys/{name}", response_model=AIKeyDeleteResponse)
async def api_delete_tool_key(name: str, current_user=Depends(require_admin)):
"""Delete a stored tool key (the environment fallback still applies)."""
key_name = name.upper()
try:
existed = _delete_tool_key(key_name)
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
logger.info(f"Tool key deleted: {key_name} (existed={existed})")
return {"status": "deleted", "key": key_name}
@app.post("/api/config/ai-keys/test", response_model=AITestResponse)
async def api_test_ai_keys(current_user=Depends(require_admin)):
"""Test which AI providers are configured.
+3
View File
@@ -20,3 +20,6 @@ psutil>=5.9
pywebpush>=2.3.0
mcp==1.9.4
sse-starlette==2.1.3
openpyxl>=3.1
python-docx>=1.1
reportlab>=4.0
+45 -3
View File
@@ -34,7 +34,7 @@ _PATTERNS = [
(re.compile(r'(?:api[_-]?key|apikey|secret|token|password|passwd|auth[_-]?token)\s*[:=]\s*[\'"]?([^\s\'"]{20,})[\'"]?', re.IGNORECASE),
lambda m: f'{m.group(0).split("=")[0].split(":")[0]}=[MASQUÉ]' if "=" in m.group(0) or ":" in m.group(0) else '[MASQUÉ]'),
# Generic long hex/base64 strings that look like secrets (40+ chars)
# Prefixed API keys (sk-..., pk-..., rk-...)
(re.compile(r'(?:sk|pk|rk)-[a-zA-Z0-9]{20,}'), '[CLÉ API MASQUÉE]'),
# AWS access keys
@@ -43,10 +43,50 @@ _PATTERNS = [
# GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_)
(re.compile(r'gh[pousr]_[a-zA-Z0-9]{36,}'), '[GITHUB_TOKEN MASQUÉ]'),
# Generic long random-looking strings (40+ hex chars)
(re.compile(r'\b[a-fA-F0-9]{40,64}\b'), '[HEX_KEY MASQUÉ]'),
]
# BUG-035: bare 40–64 char hex strings used to be redacted unconditionally,
# which mangled legitimate git commit SHAs, checksums and hashes in notes.
# They are now only redacted when a secret-ish keyword sits in the immediate
# context; hash/commit keywords explicitly exempt them.
_HEX_RE = re.compile(r'\b[a-fA-F0-9]{40,64}\b')
_SECRET_CONTEXT_RE = re.compile(
r'(?i)\b(?:secret|token|key|apikey|api[_-]?key|password|passwd|auth|bearer|'
r'credential|x-api-key|x-auth-token)\b'
)
_HASH_CONTEXT_RE = re.compile(
r'(?i)\b(?:commit|sha\d*|hash|md5|blob|git|checksum|digest|integrity|'
r'revision|rev|etag|fingerprint)\b'
)
#: How far before the hex string a keyword may appear to count as context.
_HEX_CONTEXT_WINDOW = 60
def _redact_bare_hex_secrets(text: str) -> tuple:
"""Redact 40–64 char hex strings only when a secret keyword is nearby.
Git/SHA/checksum contexts are left untouched (BUG-035).
Args:
text: Text to scan.
Returns:
(redacted_text, redaction_count) tuple.
"""
count = 0
def _replace(match: re.Match) -> str:
nonlocal count
window = text[max(0, match.start() - _HEX_CONTEXT_WINDOW):match.start()]
if _HASH_CONTEXT_RE.search(window):
return match.group(0)
if _SECRET_CONTEXT_RE.search(window):
count += 1
return '[HEX_KEY MASQUÉ]'
return match.group(0)
return _HEX_RE.sub(_replace, text), count
def redact(text: str) -> tuple:
"""Redact sensitive patterns from text.
@@ -66,6 +106,8 @@ def redact(text: str) -> tuple:
new_result, n = pattern.subn(str(replacement), result)
count += n
result = new_result
result, hex_count = _redact_bare_hex_secrets(result)
count += hex_count
if count > 0:
logger.info(f"Redacted {count} secret(s) from content")
return result, count
+9 -3
View File
@@ -44,7 +44,7 @@ def _ensure_writable(root: Path) -> None:
raise ServiceError("Vault is read-only", code="read_only", status=403)
def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
def _validate_extension(file_path: Path, *, allow_images: bool = False, allow_docs: bool = False) -> None:
"""Reject unsupported file extensions (400)."""
from backend.indexer import SUPPORTED_EXTENSIONS
@@ -53,6 +53,9 @@ def _validate_extension(file_path: Path, *, allow_images: bool = False) -> None:
if allow_images:
from backend.attachment_indexer import IMAGE_EXTENSIONS
allowed = allowed | IMAGE_EXTENSIONS
if allow_docs:
# Office documents produced by the AI tool layer (#92).
allowed = allowed | {".xlsx", ".docx"}
if ext not in allowed and file_path.name.lower() not in ("dockerfile", "makefile"):
raise ServiceError(
@@ -715,17 +718,20 @@ def save_raw_file(
content: bytes,
*,
overwrite: bool = True,
allow_docs: bool = False,
) -> dict[str, Any]:
"""Save a binary or text file to a vault (e.g. from upload / drag-and-drop).
Creates parent directories automatically and safely validates the path.
Supports supported text extensions, images and Excalidraw files.
Supports supported text extensions, images, Excalidraw files and — with
``allow_docs`` — Office documents (.xlsx/.docx) produced by the AI tools.
Args:
vault_name: Name of the vault.
path: Vault-relative path.
content: Raw bytes to write.
overwrite: When True, replace existing files (with backup).
allow_docs: Also accept .xlsx/.docx extensions (AI document tools).
Returns:
Dict with ``success``, ``vault``, ``path``, and ``size``.
@@ -733,7 +739,7 @@ def save_raw_file(
root = get_vault_root(vault_name)
_ensure_writable(root)
file_path = resolve_safe_path(root, path)
_validate_extension(file_path, allow_images=True)
_validate_extension(file_path, allow_images=True, allow_docs=allow_docs)
rel_path = _rel(root, file_path)
+3
View File
@@ -9,6 +9,9 @@ Note: ObsiGate uses implicit namespace packages (no tracked ``__init__.py``,
which ``.gitignore`` excludes via ``_*.py``), hence this explicit facade.
"""
from backend.tools import connected as _connected # noqa: F401 (registers connected-source tools)
from backend.tools import crawler as _crawler # noqa: F401 (registers the site crawler)
from backend.tools import documents as _documents # noqa: F401 (registers document tools)
from backend.tools import service as _service # noqa: F401 (registers tools)
from backend.tools import web as _web # noqa: F401 (registers web tools)
from backend.tools.context import (
+241
View File
@@ -0,0 +1,241 @@
"""Connected sources — Gitea & GitHub repositories (phase 2 #92).
The assistant can query the source-hosting platforms the project actually
uses (ObsiGate is hosted on Gitea): repositories, issues/pull requests and
repository files. Everything is READ-risk, rate-limited through the shared
registry and audited.
Configuration (environment — injected by Infisical in production, never
hard-coded):
* ``OBSIGATE_GITEA_URL`` — base URL of the self-hosted instance (e.g.
``https://git.example.net``); the ``gitea`` provider is only available when
this variable is set. Admin-controlled, so the SSRF guard does not apply
(unlike user-supplied URLs). Both the URL and the tokens can also be set
from the configuration page (stored in ``data/api_keys.json``, #103) —
the stored value takes precedence over the environment.
* ``OBSIGATE_GITEA_TOKEN`` — optional personal access token (private repos).
* ``OBSIGATE_GITHUB_TOKEN`` — optional token (raises the API rate limits and
unlocks private repositories).
Cloud drives (Google Drive / OneDrive) deliberately stay out of the core:
per the documented roadmap they are best served by an *external MCP server*
(#79) so the OAuth surface remains outside ObsiGate.
"""
from __future__ import annotations
import base64
import binascii
import logging
from typing import Any
import httpx
from backend.tools.context import ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import GitGetFileInput, GitProviderInput, GitSearchIssuesInput
from backend.tools.secrets import get_tool_key
logger = logging.getLogger("obsigate.tools.connected")
TIMEOUT = 10.0
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
MAX_FILE_BYTES = 300_000
GITHUB_API = "https://api.github.com"
def _provider_base(provider: str) -> tuple[str, str]:
"""Return (base_url, auth_header_value) for the requested provider."""
if provider == "gitea":
base = get_tool_key("OBSIGATE_GITEA_URL").rstrip("/")
if not base:
raise ToolError(
"Source Gitea non configurée (OBSIGATE_GITEA_URL absente).",
code="provider_not_configured",
)
token = get_tool_key("OBSIGATE_GITEA_TOKEN")
return base, f"token {token}" if token else ""
if provider == "github":
token = get_tool_key("OBSIGATE_GITHUB_TOKEN")
return GITHUB_API, f"Bearer {token}" if token else ""
raise ToolError(
f"Fournisseur inconnu : {provider} ('gitea' ou 'github')",
code="invalid_arguments",
)
def _headers(auth: str) -> dict[str, str]:
headers = {"User-Agent": USER_AGENT, "Accept": "application/json"}
if auth:
headers["Authorization"] = auth
return headers
def _request(method: str, url: str, auth: str, **kwargs: Any) -> httpx.Response:
try:
resp = httpx.request(
method, url, headers=_headers(auth), timeout=TIMEOUT, follow_redirects=False,
**kwargs,
)
except httpx.HTTPError as e:
logger.warning("connected source request failed %s: %s", url, e)
raise ToolError(
"Source connectée momentanément indisponible.",
code="connected_source_unavailable",
) from e
if resp.status_code in (401, 403):
raise ToolError(
"Accès refusé par la source connectée (jeton manquant ou expiré).",
code="permission_denied",
)
if resp.status_code == 404:
raise ToolError("Ressource introuvable sur la source connectée.", code="not_found")
resp.raise_for_status()
return resp
def _normalize_repo(item: dict[str, Any]) -> dict[str, Any]:
return {
"name": item.get("name") or "",
"full_name": item.get("full_name") or "",
"url": item.get("html_url") or item.get("clone_url") or "",
"description": item.get("description") or "",
"updated": item.get("updated_at") or "",
"private": bool(item.get("private", False)),
}
@tool(
name="git_list_repos",
description=(
"List repositories on the connected Gitea instance or GitHub account "
"(name, url, description, last update). Use when the user asks about "
"their code projects."
),
input_model=GitProviderInput,
risk=ToolRisk.READ,
)
def git_list_repos(ctx, params: GitProviderInput) -> dict[str, Any]:
"""Query the configured source and return normalized repositories."""
base, auth = _provider_base(params.provider)
if params.provider == "gitea":
url = base + "/api/v1/repos/search"
query: dict[str, Any] = {"limit": params.limit}
if params.repo:
query["q"] = params.repo
resp = _request("GET", url, auth, params=query)
items = resp.json().get("data") or []
else:
if params.repo:
url = GITHUB_API + f"/repos/{params.repo.strip('/')}"
items = [_request("GET", url, auth).json()]
else:
resp = _request(
"GET", GITHUB_API + "/user/repos",
auth, params={"per_page": params.limit, "sort": "updated"},
)
items = resp.json()
repos = [_normalize_repo(item) for item in items if isinstance(item, dict)]
return {"provider": params.provider, "count": len(repos), "repos": repos}
@tool(
name="git_search_issues",
description=(
"Search issues and pull requests on the connected Gitea instance or "
"GitHub (title/body keywords, optional repository scope, open/closed)."
),
input_model=GitSearchIssuesInput,
risk=ToolRisk.READ,
)
def git_search_issues(ctx, params: GitSearchIssuesInput) -> dict[str, Any]:
"""Query issues (and PRs) from the configured source."""
base, auth = _provider_base(params.provider)
state = params.state if params.state in ("open", "closed") else "open"
if params.provider == "gitea":
if params.repo:
url = base + f"/api/v1/repos/{params.repo.strip('/')}/issues"
query: dict[str, Any] = {"state": state, "limit": params.limit, "q": params.query}
resp = _request("GET", url, auth, params=query)
items = resp.json()
else:
url = base + "/api/v1/repos/issues/search"
resp = _request("GET", url, auth, params={
"q": params.query, "state": state, "limit": params.limit,
})
items = resp.json()
else:
clause = f"{params.query} is:issue is:{state}"
if params.repo:
clause += f" repo:{params.repo.strip('/')}"
resp = _request(
"GET", GITHUB_API + "/search/issues", auth,
params={"q": clause, "per_page": params.limit},
)
items = (resp.json().get("items") or [])
issues = [
{
"id": item.get("number") or item.get("id") or "",
"title": (item.get("title") or "")[:300],
"url": item.get("html_url") or "",
"state": item.get("state") or "",
"pull_request": bool(item.get("pull_request")),
}
for item in (items if isinstance(items, list) else [])
if isinstance(item, dict)
]
return {
"provider": params.provider,
"query": params.query,
"count": len(issues),
"issues": issues,
}
@tool(
name="git_get_file",
description=(
"Read a file's content from a connected Gitea or GitHub repository "
"(source code, docs, config). Text/JSON only, size-capped."
),
input_model=GitGetFileInput,
risk=ToolRisk.READ,
)
def git_get_file(ctx, params: GitGetFileInput) -> dict[str, Any]:
"""Fetch one repository file and return its decoded text content."""
base, auth = _provider_base(params.provider)
repo = params.repo.strip("/")
path = params.path.strip("/")
if not repo or not path:
raise ToolError(
"'repo' (owner/nom) et 'path' sont obligatoires", code="invalid_arguments"
)
if params.provider == "gitea":
url = base + f"/api/v1/repos/{repo}/contents/{path}"
else:
url = GITHUB_API + f"/repos/{repo}/contents/{path}"
if params.ref:
url += f"?ref={params.ref}"
resp = _request("GET", url, auth)
data = resp.json()
encoded = data.get("content") or ""
if (data.get("encoding") or "") == "base64" and encoded:
try:
content = base64.b64decode(encoded).decode("utf-8", errors="replace")
except (ValueError, binascii.Error) as e:
raise ToolError(
"Contenu du fichier illisible (encodage inattendu).",
code="file_decode_error",
) from e
else:
content = encoded
truncated = len(content) > MAX_FILE_BYTES
return {
"provider": params.provider,
"repo": repo,
"path": data.get("path") or path,
"size": data.get("size") or len(content),
"content": content[:MAX_FILE_BYTES],
"truncated": truncated,
}
+197
View File
@@ -0,0 +1,197 @@
"""Multi-page site crawl — ``crawl_site`` (phase 2 #92, WRITE + confirmation).
The assistant can digest a small public site (documentation, docs portal) and
store a Markdown summary inside a vault: one section per page, title, URL and
readable text. The crawl is bounded and same-host only:
* max 20 pages (``max_pages``), same hostname, breadth-first from the entry URL;
* SSRF guard on every URL (scheme + private-address rejection), size caps;
* no third-party crawler dependency (scrapy deliberately avoided — a bounded
httpx BFS keeps the surface small and the runtime predictable; the task is
executed as a single background-style tool run instead of a web request
pipeline).
Risk is WRITE: the digest is written into a vault, so the two-step
confirmation applies (Apply card in the UI, propose/apply over MCP).
"""
from __future__ import annotations
import logging
import re
import time
from typing import Any
from urllib.parse import urljoin, urlparse
import httpx
from backend.tools.context import ToolContext, ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import CrawlSiteInput
from backend.tools.web import (
USER_AGENT,
_assert_public_http_url,
_html_to_text,
_response_text,
)
logger = logging.getLogger("obsigate.tools.crawler")
MAX_PAGE_BYTES = 800_000
MAX_TOTAL_BYTES = 6_000_000
MAX_TEXT_PER_PAGE = 12_000
PAGE_TIMEOUT = 10.0
_LINK_RE = re.compile(r'<a[^>]*href="([^"#]+)"', re.IGNORECASE)
_TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
def _same_host(url: str, host: str) -> bool:
return (urlparse(url).hostname or "") == host
def _extract_links(raw: str, base_url: str) -> list[str]:
import html as html_lib
links: list[str] = []
for match in _LINK_RE.finditer(raw):
href = html_lib.unescape(match.group(1)).strip()
if not href or href.lower().startswith(("javascript:", "mailto:", "tel:")):
continue
absolute = urljoin(base_url, href)
if absolute.lower().endswith((".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp", ".pdf", ".zip")):
continue
links.append(absolute.split("#", 1)[0])
return links
def _fetch_page(url: str) -> tuple[str, str]:
"""Fetch one page (SSRF-guarded, manual redirects) → (title, text)."""
current = _assert_public_http_url(url)
resp = None
for _hop in range(5):
resp = httpx.get(
current,
headers={"User-Agent": USER_AGENT, "Accept": "text/html,*/*"},
timeout=PAGE_TIMEOUT,
follow_redirects=False,
)
if resp.status_code in (301, 302, 303, 307, 308):
location = resp.headers.get("location") or ""
if not location:
break
current = _assert_public_http_url(str(httpx.URL(current).join(location)))
continue
break
assert resp is not None
resp.raise_for_status()
ctype = (resp.headers.get("content-type") or "").lower()
if "html" not in ctype and "text" not in ctype:
raise ToolError(
f"Type de contenu non pris en charge: {ctype.split(';')[0] or 'inconnu'}",
code="unsupported_content_type",
)
raw = (resp.content[:MAX_PAGE_BYTES]).decode(resp.encoding or "utf-8", errors="replace")
title_match = _TITLE_RE.search(raw)
import html as html_lib
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
return title, _html_to_text(raw)[:MAX_TEXT_PER_PAGE]
@tool(
name="crawl_site",
description=(
"Crawl a small public site (same-host only, max 20 pages) starting at "
"a URL and save a Markdown digest (title, url, readable text per page) "
"into a vault. Use to capture an online documentation for offline use."
),
input_model=CrawlSiteInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def crawl_site(ctx: ToolContext, params: CrawlSiteInput) -> dict[str, Any]:
"""Bounded BFS crawl; writes the digest file and returns a summary."""
from backend.services.errors import ServiceError
from backend.services.mutations import save_raw_file
start = _assert_public_http_url(params.url.strip())
host = urlparse(start).hostname or ""
if not host:
raise ToolError("URL sans hôte", code="invalid_url")
queue: list[str] = [start]
seen: set[str] = {start}
pages: list[dict[str, Any]] = []
total_bytes = 0
failures: list[str] = []
while queue and len(pages) < params.max_pages and total_bytes < MAX_TOTAL_BYTES:
url = queue.pop(0)
try:
title, text = _fetch_page(url)
except ToolError as e:
failures.append(url)
logger.warning("crawl_site page failed %s: %s", url, e.code)
continue
except httpx.HTTPError as e:
failures.append(url)
logger.warning("crawl_site page failed %s: %s", url, e)
continue
pages.append({"url": url, "title": title, "text": text})
total_bytes += len(text)
if len(pages) >= params.max_pages:
break
try:
raw_resp = httpx.get(
url, headers={"User-Agent": USER_AGENT}, timeout=PAGE_TIMEOUT,
follow_redirects=False,
)
raw = _response_text(raw_resp)
except (httpx.HTTPError, ValueError):
continue
for link in _extract_links(raw, url):
if len(pages) + len(queue) >= params.max_pages:
break
if link in seen or not _same_host(link, host):
continue
try:
_assert_public_http_url(link)
except ToolError:
continue
seen.add(link)
queue.append(link)
if not pages:
raise ToolError(
"Aucune page n'a pu être récupérée pour ce site.",
code="crawl_failed",
)
lines = [
f"# Crawl de {host}",
"",
f"> {len(pages)} page(s) capturée(s) depuis {start} — {time.strftime('%Y-%m-%d %H:%M')}",
"",
]
for page in pages:
lines.append(f"## {page['title'] or page['url']}")
lines.append("")
lines.append(f"Source : {page['url']}")
lines.append("")
lines.append(page["text"])
lines.append("")
digest = "\n".join(lines).encode("utf-8")
try:
saved = save_raw_file(
params.vault, params.path, digest, overwrite=True, allow_docs=False
)
except ServiceError as e:
raise ToolError(e.message, code=e.code, details=e.details) from e
return {
"url": start,
"vault": params.vault,
"path": saved.get("path", params.path),
"pages": len(pages),
"failed": failures[:10],
"size": saved.get("size", len(digest)),
}
+229
View File
@@ -0,0 +1,229 @@
"""Document-production tools (phase 2 #92) — WRITE, confirmation required.
The assistant can generate real files inside a vault:
* ``create_xlsx`` — spreadsheet (openpyxl);
* ``create_docx`` — Word document (python-docx);
* ``create_csv`` — CSV (stdlib);
* ``create_pdf`` — PDF (reportlab, from markdown-ish content).
Every tool is ``WRITE`` (two-step confirm in the UI / propose-apply over MCP),
vault-scoped through ``requires_vault`` and saved via the shared mutation
service (path safety, read-only check, backup on overwrite).
"""
from __future__ import annotations
import csv as csv_lib
import io
import logging
import re
from typing import Any
from xml.sax import saxutils
from backend.services.errors import ServiceError
from backend.services.mutations import save_raw_file
from backend.tools.context import ToolContext, ToolError, ToolRisk
from backend.tools.registry import tool
from backend.tools.schemas import CsvInput, DocxInput, PdfInput, SpreadsheetInput
logger = logging.getLogger("obsigate.tools.documents")
MAX_PDF_CHARS = 200_000
MAX_ROWS = 5_000
def _save(vault: str, path: str, content: bytes, overwrite: bool) -> dict[str, Any]:
"""Shared save helper (maps ServiceError to ToolError)."""
try:
return save_raw_file(vault, path, content, overwrite=overwrite, allow_docs=True)
except ServiceError as e:
raise ToolError(e.message, code=e.code, details=e.details) from e
def _check_rows(rows: list[list[Any]]) -> None:
if not rows:
raise ToolError("Aucune ligne fournie", code="invalid_arguments")
if len(rows) > MAX_ROWS:
raise ToolError(
f"Trop de lignes ({len(rows)} > {MAX_ROWS})", code="invalid_arguments"
)
def _check_extension(path: str, expected: str) -> str:
"""Enforce the document extension; return the normalized path."""
path = (path or "").strip()
if not path.lower().endswith(expected):
raise ToolError(
f"Extension attendue : {expected}", code="invalid_arguments"
)
return path
@tool(
name="create_xlsx",
description=(
"Create an .xlsx spreadsheet in a vault from rows of cell values "
"(first row = header). Use for tables, budgets, checklists the user "
"asked to turn into an Excel file."
),
input_model=SpreadsheetInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_xlsx(ctx: ToolContext, params: SpreadsheetInput) -> dict[str, Any]:
"""Build the workbook with openpyxl and save it into the vault."""
from openpyxl import Workbook
_check_rows(params.rows)
path = _check_extension(params.path, ".xlsx")
wb = Workbook()
ws = wb.active
ws.title = params.sheet_name[:31] or "Feuille1"
for row in params.rows:
ws.append(list(row))
buffer = io.BytesIO()
wb.save(buffer)
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
@tool(
name="create_docx",
description=(
"Create a .docx Word document in a vault from an optional title and "
"ordered paragraphs. Use for letters, reports, structured drafts."
),
input_model=DocxInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_docx(ctx: ToolContext, params: DocxInput) -> dict[str, Any]:
"""Build the document with python-docx and save it into the vault."""
from docx import Document
if not params.paragraphs:
raise ToolError("Aucun paragraphe fourni", code="invalid_arguments")
path = _check_extension(params.path, ".docx")
doc = Document()
if params.title.strip():
doc.add_heading(params.title.strip(), level=1)
for paragraph in params.paragraphs:
doc.add_paragraph(paragraph)
buffer = io.BytesIO()
doc.save(buffer)
return _save(params.vault, path, buffer.getvalue(), params.overwrite)
@tool(
name="create_csv",
description=(
"Create a .csv file in a vault from rows of cell values (first row = "
"header). Use for flat data exports, simple tables."
),
input_model=CsvInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_csv(ctx: ToolContext, params: CsvInput) -> dict[str, Any]:
"""Serialize the rows and save the CSV into the vault."""
_check_rows(params.rows)
path = _check_extension(params.path, ".csv")
delimiter = params.delimiter if params.delimiter in (",", ";", "\t") else ","
buffer = io.StringIO()
writer = csv_lib.writer(buffer, delimiter=delimiter, lineterminator="\n")
writer.writerows(params.rows)
return _save(params.vault, path, buffer.getvalue().encode("utf-8"), params.overwrite)
_HEADING_RE = re.compile(r"^(#{1,6})\s+(.*)$")
def _markdown_to_flowables(content: str) -> list[tuple[str, str]]:
"""Split markdown-ish content into (style, text) blocks for reportlab."""
blocks: list[tuple[str, str]] = []
for raw_line in content.splitlines():
line = raw_line.rstrip()
if not line.strip():
continue
heading = _HEADING_RE.match(line)
if heading:
blocks.append((f"H{min(3, len(heading.group(1)))}", heading.group(2).strip()))
else:
blocks.append(("P", line.strip()))
return blocks
def _render_markdown_pdf(content: str, title: str) -> bytes | None:
"""Render markdown → HTML → PDF through the document-page pipeline.
Uses the same stack as the « Download PDF » button of the document viewer
(mistune with the table plugin + WeasyPrint print CSS), so tables, code
blocks and lists are laid out correctly. Returns ``None`` when WeasyPrint
is not importable (missing GTK on some hosts) so the caller can fall back
to the simplified reportlab renderer.
"""
try:
import mistune
from backend.pdf_export import build_pdf_html, generate_pdf
renderer = mistune.create_markdown(
escape=False,
plugins=["table", "strikethrough", "footnotes", "task_lists"],
)
html = renderer(content)
return generate_pdf(build_pdf_html(html, title), title)
except Exception as e:
# WeasyPrint loads GTK lazily: a missing native library can surface at
# import OR render time. Fall back to the simple renderer either way.
logger.warning("WeasyPrint pipeline unavailable for create_pdf: %s", e)
return None
def _render_reportlab_pdf(content: str, title: str) -> bytes:
"""Fallback renderer (no WeasyPrint): headings + paragraphs, no tables."""
from reportlab.lib.pagesizes import A4
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.platypus import Paragraph, SimpleDocTemplate, Spacer
styles = getSampleStyleSheet()
style_map = {
"P": styles["BodyText"],
"H1": styles["Heading1"],
"H2": styles["Heading2"],
"H3": styles["Heading3"],
}
buffer = io.BytesIO()
doc = SimpleDocTemplate(buffer, pagesize=A4, title=title[:200])
story: list[Any] = [Paragraph(saxutils.escape(title[:300]), styles["Title"])]
for style, line in _markdown_to_flowables(content):
story.append(Spacer(1, 4))
story.append(Paragraph(saxutils.escape(line), style_map[style]))
doc.build(story)
return buffer.getvalue()
@tool(
name="create_pdf",
description=(
"Create a .pdf document in a vault from markdown content (headings, "
"paragraphs, tables, code blocks, lists). Use for printable "
"deliverables; tables are laid out like the document-page PDF export."
),
input_model=PdfInput,
risk=ToolRisk.WRITE,
requires_vault=True,
)
def create_pdf(ctx: ToolContext, params: PdfInput) -> dict[str, Any]:
"""Render the content and save the PDF into the vault.
Primary path: mistune (tables) + WeasyPrint — identical to the viewer's
« Download PDF » export. Fallback (WeasyPrint unavailable): simplified
reportlab layout without tables.
"""
path = _check_extension(params.path, ".pdf")
content = params.content[:MAX_PDF_CHARS]
pdf_bytes = _render_markdown_pdf(content, params.title[:300])
if pdf_bytes is None:
pdf_bytes = _render_reportlab_pdf(content, params.title[:300])
return _save(params.vault, path, pdf_bytes, params.overwrite)
+8
View File
@@ -47,6 +47,14 @@ _STEP_LABELS: dict[str, tuple[str, str | None]] = {
"restore_backup": ("backup_restore", "path"),
"web_search": ("web_search", "query"),
"fetch_url": ("fetch_url", "url"),
"crawl_site": ("crawl", "url"),
"git_list_repos": ("git_repos", "provider"),
"git_search_issues": ("git_issues", "query"),
"git_get_file": ("git_file", "path"),
"create_xlsx": ("xlsx_create", "path"),
"create_docx": ("docx_create", "path"),
"create_csv": ("csv_create", "path"),
"create_pdf": ("pdf_create", "path"),
}
GENERIC_KEY = "generic"
+84
View File
@@ -251,6 +251,90 @@ class FetchUrlInput(BaseModel):
"""Fetch one public web page and return its readable text."""
url: str = Field(..., description="Absolute http(s) URL of a public page")
render: bool = Field(
False,
description="Render JavaScript with the optional Playwright worker (dynamic SPA pages)",
)
class CrawlSiteInput(BaseModel):
"""Crawl a small public site (same-host only) and save a digest into a vault."""
url: str = Field(..., description="Absolute http(s) URL where the crawl starts")
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the digest file to write (.md)")
max_pages: int = Field(5, ge=1, le=20, description="Maximum number of pages to crawl")
class GitProviderInput(BaseModel):
"""Base fields for connected-source tools (Gitea / GitHub)."""
provider: str = Field(..., description="'gitea' (OBSIGATE_GITEA_URL) or 'github'")
repo: str = Field("", description="Optional 'owner/name' repository filter")
limit: int = Field(20, ge=1, le=50, description="Maximum number of entries")
class GitSearchIssuesInput(BaseModel):
"""Search issues/pull requests on a connected Gitea or GitHub instance."""
provider: str = Field(..., description="'gitea' or 'github'")
query: str = Field(..., min_length=1, description="Search keywords")
repo: str = Field("", description="Optional 'owner/name' scope (empty = instance-wide)")
state: str = Field("open", description="'open' or 'closed'")
limit: int = Field(10, ge=1, le=20, description="Maximum number of issues")
class GitGetFileInput(BaseModel):
"""Read a file from a connected Gitea or GitHub repository."""
provider: str = Field(..., description="'gitea' or 'github'")
repo: str = Field(..., description="'owner/name' repository")
path: str = Field(..., description="Repository-relative file path")
ref: str = Field("", description="Optional branch/tag/commit (empty = default branch)")
class SpreadsheetInput(BaseModel):
"""Create an .xlsx spreadsheet in a vault from rows of cells."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.xlsx)")
rows: list[list[str | int | float | bool | None]] = Field(
..., description="Rows of cell values (first row = header)"
)
sheet_name: str = Field("Feuille1", description="Worksheet name")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class DocxInput(BaseModel):
"""Create a .docx Word document in a vault from paragraphs."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.docx)")
title: str = Field("", description="Optional document title (heading 1)")
paragraphs: list[str] = Field(..., description="Paragraph texts, in order")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class CsvInput(BaseModel):
"""Create a .csv file in a vault from rows of cells."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.csv)")
rows: list[list[str | int | float | bool | None]] = Field(
..., description="Rows of cell values (first row = header)"
)
delimiter: str = Field(",", description="Field separator (',' ';' '\\t')")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class PdfInput(BaseModel):
"""Create a .pdf document in a vault from markdown-ish content."""
vault: str = Field(..., description="Vault name")
path: str = Field(..., description="Vault-relative path of the file to write (.pdf)")
title: str = Field("Document", description="Document title")
content: str = Field(..., description="Content (headings with #/##, then paragraphs)")
overwrite: bool = Field(True, description="Replace an existing file (with backup)")
class ToolResult(BaseModel):
+109
View File
@@ -0,0 +1,109 @@
"""Tool-layer secrets — user-configured tokens & API keys (#103).
The connected-source (Gitea / GitHub) and keyed web-search (Tavily, Brave,
SerpAPI, Exa) tools read their credentials through this module instead of
``os.environ`` directly. The value comes from the store the user edits in the
configuration page (``data/api_keys.json`` — the same file the AI provider
keys use) first, then falls back to the environment (Infisical-injected in
production). Nothing is ever hard-coded and no tool result carries a secret
(the registry redacts payloads).
Allowed names are whitelisted: only the variables below can be stored or
deleted from the configuration page.
"""
from __future__ import annotations
import json
import logging
import os
from pathlib import Path
logger = logging.getLogger("obsigate.tools.secrets")
# Whitelisted configuration names (config page « Sources connectées & recherche »).
TOOL_KEY_NAMES: tuple[str, ...] = (
"OBSIGATE_TAVILY_API_KEY",
"OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY",
"OBSIGATE_EXA_API_KEY",
"OBSIGATE_GITEA_URL",
"OBSIGATE_GITEA_TOKEN",
"OBSIGATE_GITHUB_TOKEN",
)
_SECRET_MARKERS = ("API_KEY", "TOKEN")
def _keys_file() -> Path:
base = os.environ.get("OBSIGATE_DATA_DIR", "data")
return Path(base) / "api_keys.json"
def _read_keys() -> dict:
path = _keys_file()
if not path.exists():
return {}
try:
data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, ValueError) as e:
logger.warning("tool key store unreadable (%s): %s", path, e)
return {}
return data if isinstance(data, dict) else {}
def _write_keys(data: dict) -> None:
path = _keys_file()
path.parent.mkdir(parents=True, exist_ok=True)
tmp = path.with_suffix(".tmp")
tmp.write_text(json.dumps(data, indent=2), encoding="utf-8")
tmp.replace(path)
def is_secret_name(name: str) -> bool:
"""True for API keys / tokens (masked in API responses); URLs are clear."""
return any(marker in name for marker in _SECRET_MARKERS)
def mask_value(name: str, value: str) -> str:
"""Mask a secret for display; non-secret values (URLs) are returned as-is."""
if not value:
return ""
if not is_secret_name(name):
return value
return value[:4] + "..." + value[-4:] if len(value) > 8 else "***"
def get_tool_key(name: str) -> str:
"""Stored (configuration page) value first, then environment fallback."""
if name not in TOOL_KEY_NAMES:
return os.environ.get(name, "").strip()
stored = _read_keys().get(name)
if isinstance(stored, str) and stored.strip():
return stored.strip()
return os.environ.get(name, "").strip()
def set_tool_key(name: str, value: str) -> None:
"""Persist one whitelisted key into the store (admin configuration page)."""
if name not in TOOL_KEY_NAMES:
raise ValueError(f"Clé non prise en charge: {name}")
value = (value or "").strip()
keys = _read_keys()
if value:
keys[name] = value
else:
keys.pop(name, None)
_write_keys(keys)
def delete_tool_key(name: str) -> bool:
"""Remove one key from the store; return True when it existed."""
if name not in TOOL_KEY_NAMES:
raise ValueError(f"Clé non prise en charge: {name}")
keys = _read_keys()
if name in keys:
del keys[name]
_write_keys(keys)
return True
return False
+181 -4
View File
@@ -8,6 +8,16 @@ Phase 1 of the documented web-toolset roadmap:
answering « je n'ai pas accès à internet ».
* ``fetch_url`` — retrieve a public web page and return readable text.
Phase 2 (#92) additions:
* keyed providers — Tavily, Brave Search, SerpAPI and Exa are used first when
their API key is configured (env, injected by Infisical in production);
* SQLite cache — search/fetch results are cached with a TTL
(:mod:`backend.tools.webcache`);
* retry with backoff — transient network errors get one extra attempt;
* dynamic rendering — ``fetch_url(render=True)`` uses an isolated Playwright
worker (optional dependency, graceful degradation).
All are READ-risk tools (no confirmation), rate-limited through the shared
registry, SSRF-guarded (scheme + private-address rejection), and size-capped.
@@ -16,6 +26,14 @@ Configuration (environment):
* ``OBSIGATE_WEB_TIMEOUT`` — seconds, default 10
* ``OBSIGATE_WEB_FALLBACK`` — ``0``/``false`` disables the keyless HTML
fallbacks (SearXNG only), default enabled
* ``OBSIGATE_TAVILY_API_KEY`` / ``OBSIGATE_BRAVE_API_KEY`` /
``OBSIGATE_SERPAPI_API_KEY`` / ``OBSIGATE_EXA_API_KEY`` — optional keyed
providers, tried before SearXNG when set
* ``OBSIGATE_WEB_PROVIDERS`` — optional comma-separated provider order
(e.g. ``brave,searxng``); keyed providers without a key are skipped
* ``OBSIGATE_WEB_RETRY`` — extra attempts for transient network errors
(default 1)
* ``OBSIGATE_WEB_CACHE_TTL`` — cache TTL seconds, ``0`` disables (default 900)
"""
from __future__ import annotations
@@ -28,15 +46,18 @@ import logging
import os
import re
import socket
import time
from collections.abc import Callable
from typing import Any
from urllib.parse import parse_qs, urlparse
import httpx
from backend.tools import webcache
from backend.tools.context import ToolError, ToolRisk, ToolScope
from backend.tools.registry import tool
from backend.tools.schemas import FetchUrlInput, WebSearchInput
from backend.tools.secrets import get_tool_key
logger = logging.getLogger("obsigate.tools.web")
@@ -48,6 +69,7 @@ WEB_FALLBACK_ENABLED = os.environ.get("OBSIGATE_WEB_FALLBACK", "1").strip().lowe
"no",
"off",
}
WEB_RETRY_ATTEMPTS = int(os.environ.get("OBSIGATE_WEB_RETRY", "1"))
USER_AGENT = "ObsiGateAssistant/1.0 (+self-hosted vault AI)"
# Search engines reject non-browser agents on their public HTML endpoints.
BROWSER_UA = (
@@ -163,6 +185,115 @@ def _result(
}
def _with_retry(call: Callable[[], Any]) -> Any:
"""Run *call* with one extra attempt on transient network errors.
House-made backoff (the roadmap's « tenacity ou boucle maison »): DNS
blips and rate-limit hiccups are the common failure mode, and a single
retry keeps the fallback chain from being consumed too early.
"""
for attempt in range(1 + max(0, WEB_RETRY_ATTEMPTS)):
try:
return call()
except httpx.TransportError:
if attempt >= max(0, WEB_RETRY_ATTEMPTS):
raise
time.sleep(0.2 * (attempt + 1))
raise RuntimeError("unreachable") # pragma: no cover
def _env_key(name: str) -> str:
"""Read an API key: configuration-page store first, then environment."""
return get_tool_key(name)
def _search_tavily(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Tavily Search API (agent-oriented results, key required)."""
resp = httpx.post(
"https://api.tavily.com/search",
json={
"api_key": _env_key("OBSIGATE_TAVILY_API_KEY"),
"query": query,
"max_results": params.max_results,
"search_depth": "basic",
"include_answer": False,
},
headers={"User-Agent": USER_AGENT},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", item.get("content") or "")
for item in (data.get("results") or [])
], []
def _search_brave(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Brave Search API (key required)."""
resp = httpx.get(
"https://api.search.brave.com/res/v1/web/search",
params={"q": query, "count": params.max_results, "safesearch": "moderate"},
headers={
"X-Subscription-Id": _env_key("OBSIGATE_BRAVE_API_KEY"),
"Accept": "application/json",
"User-Agent": USER_AGENT,
},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", item.get("description") or "")
for item in ((data.get("web") or {}).get("results") or [])
], []
def _search_serpapi(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""SerpAPI (Google SERP, key required)."""
resp = httpx.get(
"https://serpapi.com/search",
params={"q": query, "api_key": _env_key("OBSIGATE_SERPAPI_API_KEY"),
"num": params.max_results},
headers={"User-Agent": USER_AGENT},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("link") or "", item.get("snippet") or "")
for item in (data.get("organic_results") or [])
], []
def _search_exa(query: str, params: WebSearchInput) -> tuple[list[dict[str, Any]], list[str]]:
"""Exa neural search (key required)."""
resp = httpx.post(
"https://api.exa.ai/search",
json={"query": query, "numResults": params.max_results},
headers={
"x-api-key": _env_key("OBSIGATE_EXA_API_KEY"),
"User-Agent": USER_AGENT,
},
timeout=WEB_TIMEOUT,
)
resp.raise_for_status()
data = resp.json()
return [
_result(item.get("title") or "", item.get("url") or "", (item.get("text") or "")[:600])
for item in (data.get("results") or [])
], []
# Keyed providers: name -> (implementation, API key env var)
_KEYED_PROVIDERS: dict[str, tuple[_Provider, str]] = {
"tavily": (_search_tavily, "OBSIGATE_TAVILY_API_KEY"),
"brave": (_search_brave, "OBSIGATE_BRAVE_API_KEY"),
"serpapi": (_search_serpapi, "OBSIGATE_SERPAPI_API_KEY"),
"exa": (_search_exa, "OBSIGATE_EXA_API_KEY"),
}
def _search_searxng(
query: str, params: WebSearchInput
) -> tuple[list[dict[str, Any]], list[str]]:
@@ -288,8 +419,22 @@ _Provider = Callable[[str, WebSearchInput], "tuple[list[dict[str, Any]], list[st
def _provider_chain() -> list[tuple[str, _Provider]]:
"""Ordered providers: self-hosted meta-search first, then keyless fallbacks."""
chain: list[tuple[str, _Provider]] = [("searxng", _search_searxng)]
"""Ordered providers: keyed APIs first, then self-hosted, then keyless.
``OBSIGATE_WEB_PROVIDERS`` (comma-separated) overrides the default order;
unknown names are ignored and keyed providers without their key are skipped.
"""
chain: list[tuple[str, _Provider]] = []
configured = [
name.strip().lower()
for name in os.environ.get("OBSIGATE_WEB_PROVIDERS", "").split(",")
if name.strip()
]
for name in configured or list(_KEYED_PROVIDERS):
entry = _KEYED_PROVIDERS.get(name)
if entry and _env_key(entry[1]):
chain.append((name, entry[0]))
chain.append(("searxng", _search_searxng))
if WEB_FALLBACK_ENABLED:
chain.append(("duckduckgo", _search_duckduckgo))
chain.append(("bing", _search_bing))
@@ -313,6 +458,17 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
if not query:
raise ToolError("Requête vide", code="invalid_arguments")
key = webcache.cache_key("search", {
"q": query,
"max_results": params.max_results,
"category": params.category,
"language": params.language,
"page": params.page,
})
cached = webcache.cache_get(key)
if cached is not None:
return {**cached, "cached": True}
attempts: list[str] = []
unresponsive: list[str] = []
reachable = False
@@ -320,8 +476,12 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
for name, provider in _provider_chain():
attempts.append(name)
def _attempt(p: _Provider = provider) -> tuple[list[dict[str, Any]], list[str]]:
return p(query, params)
try:
results, engines = provider(query, params)
results, engines = _with_retry(_attempt)
except (httpx.HTTPError, ValueError, AttributeError) as e:
logger.warning("web_search provider %s failed: %s", name, e)
last_error = e
@@ -338,6 +498,7 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
}
if unresponsive:
payload["unresponsive_engines"] = unresponsive[:8]
webcache.cache_set(key, payload)
return payload
if not reachable:
@@ -378,6 +539,20 @@ def web_search(ctx, params: WebSearchInput) -> dict[str, Any]:
def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
"""Retrieve one page, guard against SSRF, and extract its text."""
url = _assert_public_http_url(params.url.strip())
key = webcache.cache_key("fetch", {"url": url, "render": params.render})
cached = webcache.cache_get(key)
if cached is not None:
return {**cached, "cached": True}
if params.render:
# Dynamic pages (SPA/React): delegated to the isolated Playwright
# worker; the browser dependency stays optional (graceful error).
from backend.tools.webrender import render_page
payload = render_page(url)
webcache.cache_set(key, payload)
return payload
try:
# Follow redirects manually so every hop is re-checked against the
# private-address SSRF guard (a public page can redirect to 127.0.0.1).
@@ -414,10 +589,12 @@ def fetch_url(ctx, params: FetchUrlInput) -> dict[str, Any]:
title_match = re.search(r"<title[^>]*>(.*?)</title>", raw, re.IGNORECASE | re.DOTALL)
title = html_lib.unescape(title_match.group(1)).strip()[:300] if title_match else ""
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
return {
payload = {
"url": str(resp.url),
"status": resp.status_code,
"title": title,
"text": text,
"truncated": len(raw) > MAX_TEXT_CHARS,
}
webcache.cache_set(key, payload)
return payload
+138
View File
@@ -0,0 +1,138 @@
"""SQLite cache for web tool results (search results, fetched pages).
Phase 2 of the web-toolset roadmap (« Transverse »): repeated web searches and
page fetches (common in agent loops, where the model re-reads a source) must
not hammer the providers. Results are cached in a dedicated SQLite table with
a TTL; the cache is best-effort — any error silently disables it so a broken
database file never takes the assistant down.
Configuration (environment):
* ``OBSIGATE_DATA_DIR`` — base data directory (default ``data``)
* ``OBSIGATE_WEB_CACHE_PATH`` — explicit cache file override
* ``OBSIGATE_WEB_CACHE_TTL`` — seconds, ``0`` disables the cache (default 900)
"""
from __future__ import annotations
import hashlib
import json
import logging
import os
import sqlite3
import threading
import time
from pathlib import Path
from typing import Any
logger = logging.getLogger("obsigate.tools.webcache")
DEFAULT_TTL_SECONDS = 900
_schema_ready = False
_write_lock = threading.Lock()
def ttl_seconds() -> float:
"""Configured TTL in seconds (``0`` = cache disabled)."""
return float(os.environ.get("OBSIGATE_WEB_CACHE_TTL", str(DEFAULT_TTL_SECONDS)))
def _cache_path() -> Path:
override = os.environ.get("OBSIGATE_WEB_CACHE_PATH", "").strip()
if override:
return Path(override)
return Path(os.environ.get("OBSIGATE_DATA_DIR", "data")) / "web_cache.sqlite3"
def _connect() -> sqlite3.Connection:
"""Open (and lazily create) the cache database."""
global _schema_ready
path = _cache_path()
path.parent.mkdir(parents=True, exist_ok=True)
conn = sqlite3.connect(path, timeout=5, check_same_thread=False)
if not _schema_ready:
conn.execute(
"CREATE TABLE IF NOT EXISTS web_cache ("
"key TEXT PRIMARY KEY, value TEXT NOT NULL, created REAL NOT NULL)"
)
conn.commit()
_schema_ready = True
return conn
def cache_key(prefix: str, payload: dict[str, Any]) -> str:
"""Deterministic cache key from a prefix and the normalized arguments."""
raw = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str)
digest = hashlib.sha256(raw.encode("utf-8")).hexdigest()[:32]
return f"{prefix}:{digest}"
def cache_get(key: str) -> Any | None:
"""Return the cached payload for *key*, or ``None`` (miss/expiry/disabled)."""
if ttl_seconds() <= 0:
return None
try:
conn = _connect()
row = conn.execute(
"SELECT value, created FROM web_cache WHERE key = ?", (key,)
).fetchone()
conn.close()
except sqlite3.Error as e:
logger.warning("web cache read failed (%s): %s", key, e)
return None
if row is None:
return None
value, created = row
if time.time() - float(created) > ttl_seconds():
return None
try:
return json.loads(value)
except (ValueError, TypeError):
return None
def cache_set(key: str, value: Any) -> None:
"""Store *value* under *key* (best effort, never raises)."""
if ttl_seconds() <= 0:
return
try:
with _write_lock:
conn = _connect()
conn.execute(
"INSERT INTO web_cache (key, value, created) VALUES (?, ?, ?) "
"ON CONFLICT(key) DO UPDATE SET value = excluded.value, created = excluded.created",
(key, json.dumps(value, ensure_ascii=False, default=str), time.time()),
)
conn.commit()
conn.close()
except sqlite3.Error as e:
logger.warning("web cache write failed (%s): %s", key, e)
def purge_expired() -> int:
"""Delete expired rows; return the number of removed entries (maintenance)."""
try:
conn = _connect()
cursor = conn.execute(
"DELETE FROM web_cache WHERE created < ?", (time.time() - ttl_seconds(),)
)
conn.commit()
deleted = cursor.rowcount
conn.close()
return int(deleted)
except sqlite3.Error as e:
logger.warning("web cache purge failed: %s", e)
return 0
def clear_cache() -> int:
"""Drop every cached entry (tests / admin); returns the number of rows."""
try:
conn = _connect()
cursor = conn.execute("DELETE FROM web_cache")
conn.commit()
deleted = cursor.rowcount
conn.close()
return int(deleted)
except sqlite3.Error as e:
logger.warning("web cache clear failed: %s", e)
return 0
+100
View File
@@ -0,0 +1,100 @@
"""Dynamic page rendering (Playwright) — ``fetch_url(render=True)``.
Static pages are fetched with httpx inside :mod:`backend.tools.web`. Dynamic
pages (SPA/React, JS-loaded content) need a real browser engine; this module
runs one Playwright call inside a dedicated worker thread so browser
crashes/timeouts never take over the tool layer, and the heavyweight
dependency stays optional:
* not installed → ``ToolError(code="playwright_unavailable")`` with a clear
message (the assistant explains the limitation instead of hanging);
* installed → ``pip install playwright && playwright install chromium``.
The SSRF guard (scheme + private-address rejection) is applied before the
browser navigates. Note: unlike the httpx path, internal redirects performed
by the browser engine are not re-checked hop by hop.
"""
from __future__ import annotations
import html as html_lib
import logging
import re
from concurrent.futures import ThreadPoolExecutor
from typing import Any
from backend.tools.context import ToolError
from backend.tools.web import (
MAX_TEXT_CHARS,
USER_AGENT,
_assert_public_http_url,
_html_to_text,
)
logger = logging.getLogger("obsigate.tools.webrender")
# One worker: browser automation is serialized on purpose (one Chromium at a
# time keeps memory predictable on small hosts).
_executor = ThreadPoolExecutor(max_workers=1, thread_name_prefix="obsigate-playwright")
GOTO_TIMEOUT_MS = 20_000
def _playwright_available() -> bool:
try:
import playwright # noqa: F401
except ImportError:
return False
return True
def _render_in_worker(url: str) -> dict[str, Any]:
"""Synchronous Playwright render — runs in the dedicated worker thread."""
from playwright.sync_api import sync_playwright
status = 0
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
page = browser.new_page(user_agent=USER_AGENT)
response = page.goto(url, wait_until="networkidle", timeout=GOTO_TIMEOUT_MS)
if response is not None:
status = response.status
raw = page.content()
title = html_lib.unescape(page.title() or "").strip()
text = _html_to_text(raw)[:MAX_TEXT_CHARS]
finally:
browser.close()
title = re.sub(r"\s+", " ", title)[:300]
return {
"url": url,
"status": status,
"title": title,
"text": text,
"rendered": True,
"truncated": len(raw) > MAX_TEXT_CHARS,
}
def render_page(url: str) -> dict[str, Any]:
"""Render *url* (JavaScript included) and return readable text.
Raises:
ToolError: ``playwright_unavailable`` when the optional dependency is
missing, ``render_unavailable`` when the render itself failed.
"""
_assert_public_http_url(url)
if not _playwright_available():
raise ToolError(
"Rendu dynamique indisponible : Playwright n'est pas installé "
"(pip install playwright && playwright install chromium).",
code="playwright_unavailable",
)
try:
return _executor.submit(_render_in_worker, url).result(timeout=GOTO_TIMEOUT_MS / 1000 + 40)
except ToolError:
raise
except Exception as e:
logger.warning("render_page failed for %s: %s", url, e)
raise ToolError(
"Le rendu dynamique de la page a échoué.", code="render_unavailable"
) from e
+1 -1
View File
@@ -2626,7 +2626,7 @@ dependencies = [
[[package]]
name = "obsigate-desktop"
version = "2.8.2"
version = "2.11.3"
dependencies = [
"chrono",
"env_logger",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "obsigate-desktop"
version = "2.8.2"
version = "2.11.3"
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
authors = ["Bruno Charest"]
edition = "2021"
+1 -1
View File
@@ -1,7 +1,7 @@
{
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
"productName": "ObsiGate",
"version": "2.8.2",
"version": "2.11.3",
"identifier": "com.obsigate.desktop",
"build": {
"frontendDist": "../frontend",
+16 -6
View File
@@ -144,12 +144,12 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-032* | [🟡 IMPORTANT] Indexation : symlinks suivis (contenu hors vault indexé) + scan initial coûteux | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/indexer.py` | Placer un symlink dans le vault vers un dossier externe puis relancer l'index | `_scan_vault` réécrit avec `os.walk(followlinks=False)` + refus des symlinks sortant de la racine ; test `TestSymlinkIndexing` | Scan incrémental/index persistant : voir #86 (phase 3) |
| *BUG-033* | [🟡 IMPORTANT] Recherche classique et tool IA `search_fulltext` en O(N) sans inverted index | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/search.py`, `backend/tools/service.py` | `GET /api/search` sur un vault de 50 000 fichiers | `search()` récupère les candidats via l'inverted index (intersection des termes + expansion de préfixes), repli sur le scan pendant la construction | `search_fulltext` en bénéficie automatiquement |
| *BUG-034* | [🟡 IMPORTANT] CSP affaiblie (`'unsafe-inline'` + CDN distants) et token d'accès en sessionStorage | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/main.py`, `frontend/js/auth.js`, `frontend/js/admin.js`, `frontend/js/sync.js` | Inspecter les en-têtes CSP ; lire sessionStorage en console | Token en mémoire + cookie HttpOnly (plus de `sessionStorage`) ; CSP durcie (`object-src 'none'`, `base-uri`, `form-action`, `frame-ancestors`). *Reste : migration nonce* | `'unsafe-inline'` conservé tant que les gestionnaires inline n'ont pas été convertis (résidu documenté) |
| *BUG-035* | [🔵 MINEUR] `secret_redactor` : faux positifs sur les hashs hex (git, SHA) | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/secret_redactor.py` | Lire une note contenant un commit git (40 caractères hexadécimaux) | Restreindre le périmètre de détection (contexte clé/token) + whitelist | Contenus mutilés dans les lectures et réponses IA |
| *BUG-036* | [🔵 MINEUR] Collab WebSocket : token en query string | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/collab.py` | Observer l'URL du websocket dans le trafic réseau | Passer le token en header / étape d'authentification initiale ; borner la taille des messages | Jeton visible dans les logs/proxys |
| *BUG-037* | [🔵 MINEUR] Compte « anonymous » administrateur si auth désactivée | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/middleware.py` | Démarrer avec l'authentification désactivée | Avertissement explicite au démarrage + refus de déploiement public sans auth | Comportement par conception mais risqué si mal configuré |
| *BUG-038* | [🔵 MINEUR] Argon2 à 64 MB par vérification : risque d'épuisement mémoire | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/password.py:8` | Lancer de nombreux `POST /api/auth/login` simultanés | Recalibrer (~19 MB, t=2, p=1, norme OWASP actuelle) + maintien du rate-limit | DoS mémoire possible sur les petites instances |
| *BUG-039* | [🔵 MINEUR] Enumération de comptes : 429 (verrouillé) vs 401 (inconnu) | 🔴 ouvert | P3 | 🔐 sécurité | IA | `backend/auth/router.py:120` | Tenter un login sur un compte verrouillé puis un nom inconnu | Répondre 401 uniforme avec un timing équivalent | Le statut HTTP distingue l'existence d'un compte |
| *BUG-040* | [🔵 MINEUR] Extraction PDF intégrale (100 ko) au scan de démarrage | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/indexer.py:465` | Démarrer sur un vault contenant de nombreux PDF | Analyser les PDF en tâche de fond / à la demande (lazy) | Ralentit fortement le démarrage et le rebuild d'index |
| *BUG-035* | [🔵 MINEUR] `secret_redactor` : faux positifs sur les hashs hex (git, SHA) | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/secret_redactor.py` | Lire une note contenant un commit git (40 caractères hexadécimaux) | Masquage hex conditionné au contexte (`_redact_bare_hex_secrets`) : secret exigé dans les 60 caractères précédents, exemption explicite pour `commit`/`sha*`/`hash`/`checksum`/`git`/`etag`. Tests : `tests/test_api_main.py::TestSecretRedactor` (+4) | Contenus mutilés dans les lectures et réponses IA |
| *BUG-036* | [🔵 MINEUR] Collab WebSocket : token en query string | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/collab.py` | Observer l'URL du websocket dans le trafic réseau | `authenticate_websocket` ne lit plus `?token=` : cookie HttpOnly `access_token` uniquement ; rejet des trames > `MAX_MESSAGE_CHARS` (16 Mio) avant analyse. Tests : `tests/test_collab.py` (+3) | Jeton visible dans les logs/proxys |
| *BUG-037* | [🔵 MINEUR] Compte « anonymous » administrateur si auth désactivée | 🟢 corrigé | P2 | 🔐 sécurité | IA | `backend/auth/middleware.py`, `backend/main.py` | Démarrer avec l'authentification désactivée | `_guard_insecure_auth()` : avertissement explicite + refus de démarrage sur bind non-loopback sans `OBSIGATE_ALLOW_INSECURE=true`. Tests : `tests/test_auth.py::TestInsecureAuthGuard` (+6) | Comportement par conception mais risqué si mal configuré |
| *BUG-038* | [🔵 MINEUR] Argon2 à 64 MB par vérification : risque d'épuisement mémoire | 🟢 corrigé | P2 | 🔐 sécurité | IA | `backend/auth/password.py` | Lancer de nombreux `POST /api/auth/login` simultanés | Recalibré à `m=19456 Kio (19 Mio), t=2, p=1` (OWASP) ; anciens hachages valides + rehash auto. Test : `tests/test_auth.py::TestPasswordHashing::test_argon2_memory_recalibrated` | DoS mémoire possible sur les petites instances |
| *BUG-039* | [🔵 MINEUR] Enumération de comptes : 429 (verrouillé) vs 401 (inconnu) | 🟢 corrigé | P3 | 🔐 sécurité | IA | `backend/auth/router.py` | Tenter un login sur un compte verrouillé puis un nom inconnu | Login uniforme : inconnu / désactivé / verrouillé / rate-limit par compte → `401 Identifiants invalides` + hachage factice (timing équivalent) ; seul le rate-limit IP reste `429`. Tests : `tests/test_auth_api.py` (+3) | Le statut HTTP distinguait l'existence d'un compte |
| *BUG-040* | [🔵 MINEUR] Extraction PDF intégrale (100 ko) au scan de démarrage | 🟢 corrigé | P2 | ⚙️ backend | IA | `backend/indexer.py`, `backend/main.py` | Démarrer sur un vault contenant de nombreux PDF | `_scan_vault` ne lit que les métadonnées ; `enrich_pdf_texts()` extrait le texte après l'index (démarrage) et après chaque réindexation. Tests : `tests/test_pdf.py` (+3) | Ralentit fortement le démarrage et le rebuild d'index |
| *BUG-041* | [🟡 IMPORTANT] Assistant IA : échec sur un répertoire vide (« Aucun fichier markdown trouvé dans ce dossier ») au lieu de répondre | 🟢 corrigé | P1 | 📱 frontend + ⚙️ backend | IA | `backend/bookslm_routes.py`, `backend/bookslm.py`, `frontend/js/bookslm.js` | Ouvrir l'assistant sur un dossier vide puis envoyer une question | `_resolve_system_prompt` dégrade vers le prompt Général + bloc « Dossier vide » (plus de 404) ; contexte applicatif `app_context` enrichi (documents ouverts, répertoire, recherche, fichiers récents) | Le 404 bloquait toute la requête. Feature #88, fiche `docs/features/ai-app-context.md`. Tests : `tests/test_bookslm.py` (+3), `tests/frontend/ai.test.mjs` |
| *BUG-042* | [🟡 IMPORTANT] Assistant IA : liens de fichiers non fiables (« File not found: ») — pas de règle déterministe nom / dossier / chemin | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js` | Cliquer les liens de fichiers/dossiers dans une réponse de l'assistant (noms avec espaces et/ou accents, chemin préfixé par le nom du vault) | `_classifyPath` distingue `name` (copie presse-papiers) / `dir` (révélation arborescence) / `file` (ouverture) ; `_activatePath()` résout le chemin contre l'index du vault (exact → suffixe → basename unique) avant d'agir ; espaces + accents pris en charge (classes Unicode `\p{L}\p{N}\p{M}`, comparaison normalisée NFC, markdown `<…>`/`%20`, code inline, mentions brutes confirmées par l'index) ; `_splitVaultPrefix` retire un préfixe `Vault/…` et ouvre dans ce vault (`_fetchPathsForVault`) | Les liens morts ouvraient un fichier inexistant. Feature #88. Tests : `tests/frontend/ai.test.mjs` (+11) |
| *BUG-043* | [🟡 IMPORTANT] Assistant IA : la liste des fournisseurs de la barre latérale ne suit pas les ajouts/retraits de clés API dans la configuration du projet | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/ai.js`, `frontend/js/bookslm.js`, `frontend/js/config.js` | Ajouter (ou supprimer) une clé de fournisseur AI dans la configuration puis observer le menu Fournisseur de l'assistant sans recharger la page | Le picker lit `/api/ai/status` **une seule fois**, à sa construction, et le panneau de l'assistant est un singleton monté pour toute la session → liste figée. Nouveau `refreshAIPickers()` (exporté par `ai.js`) qui reconstruit chaque picker monté dans son emplacement `.ai-picker-slot` (conservé même sans fournisseur configuré, donc un premier fournisseur s'y monte aussi) ; appelé après `saveAIKeys()` et `deleteAIKey()` (`config.js`) ; une sélection dont le fournisseur n'est plus configuré est purgée de `obsigate_ai_picker` (retour au défaut + modèle effacé au lieu d'un nom fantôme) | Il fallait recharger la page pour voir un nouveau fournisseur (ou en voir disparaître un). Feature #82. Tests : `tests/frontend/ai.test.mjs` (+4) |
@@ -165,6 +165,10 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-053* | [🟡 IMPORTANT] Assistant IA (mode agent) : le fichier demandé n'est pas créé — le modèle émet un bloc texte `obsigate-action` au lieu d'appeler l'outil `create_file` | 🟢 corrigé | P1 | ⚙️ backend + 🤖 ia | IA | `backend/bookslm.py`, `backend/bookslm_routes.py`, `tests/test_bookslm.py` | Mode agent, contexte Général (ou dossier vide) : « créer le fichier TestVault/sport/… avec le tableau des 84 matchs » → réponse avec un bloc ```obsigate-action``` tronqué, aucun fichier | `backend/bookslm.py` : protocole d'action scindé — `GENERAL_ACTION_TOOL_PROTOCOL` (outils natifs, interdiction des blocs `obsigate-action`) utilisé quand `agent=True`, protocole texte conservé pour le chat classique ; `backend/bookslm_routes.py` : `_resolve_system_prompt(..., agent=True)` depuis l'endpoint agent + règle « Mode agent » pour les prompts dossier/documents, `max_tokens` agent 4096 → 8192 (contenu de fichier complet) | Cause : le prompt Général enseignait encore le protocole texte alors que l'agent dispose du function calling. Tests : `tests/test_bookslm.py` (+3) |
| *BUG-054* | [🟡 IMPORTANT] Éditeur « Editer » : le bouton Sauvegarder reste bloqué sur le spinner de chargement (retour au crochet uniquement après un refresh complet) | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/utils.js` | Ouvrir un fichier → Editer → cliquer Sauvegarder (ou Ctrl+S) ; rouvrir l'éditeur : le bouton reste un spinner désactivé | Nouveau helper `resetSaveButton()` (crochet `&#10003;` + `disabled=false` + styles en ligne nettoyés) appelé à l'ouverture (`openEditor`), à la fermeture (`closeEditor`) et en cas d'échec (`saveFile`). Tests : `tests/frontend/editor-inline.test.mjs` (+4) | Le nœud `#editor-save` est partagé entre sessions : l'état « spinner + désactivé » posé par une sauvegarde manuelle n'était jamais remis à zéro (succès → fermeture puis réouverture, Forge, ou échec réseau dans le `catch`). Seul un rechargement de `index.html` restaurait le crochet |
| *BUG-055* | [🟡 IMPORTANT] Éditeur Forge : l'autocomplétion (Tab) ajoute des espaces parasites, l'effacement détruit le mot complété et la complétion fantôme est illisible | 🟢 corrigé | P1 | 📱 frontend + ⚙️ backend | IA | `frontend/editor-poc.html`, `frontend/js/autocomplete.js`, `backend/ai.py`, `.gitea/workflows/ci.yml`, `tests/frontend/forge-completion.test.mjs` (nouveau) | Forge : taper un mot, puis Tab pour compléter ; un espace (voire deux) s'insère avant le mot complété, et le retour arrière efface l'ajout. La prédiction IA s'affichait décalée (texte miroir du document entier) | **Cause** : trois gestionnaires `keydown` Tab indépendants s'exécutaient tous — l'indentation (`insertAtCursor(' ')`) s'ajoutait à la complétion de mot et à l'acceptation du ghost. **Correctif** : gestion **unifiée** de Tab (`liste ouverte > ghost > mot du document > indentation`, une seule action), helpers purs partagés (`getWordFragment`, `findWordCompletions`, `normalizeGhost`, `chooseTabAction`) dans `autocomplete.js`, liste déroulante si plusieurs candidats, dropdown positionné au curseur, ghost **positionné au curseur** (fini le miroir du document, nettoyé au déplacement/scroll), complétion de mot sans espace garanti (`normalizeGhost` tronque au premier espace) et prompt `/api/ai/inline-complete` simplifié. Tests : `tests/frontend/forge-completion.test.mjs` (28). | Cause du bug : l'indentation Tab n'était pas conditionnée à l'absence de suggestion. Le ghost re-rendait tout le texte transparent + prédiction, d'où l'impression d'espaces et les erreurs d'effacement |
| *BUG-056* | [🟡 IMPORTANT] Éditeur Forge en plein écran : l'Assistant IA s'ouvre en arrière-plan et reste invisible | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs` | Forge : passer en plein écran puis cliquer le bouton « Assistant IA » (ou `Ctrl+J`) — le panneau s'ouvre dans le document parent, masqué par l'iframe plein écran | Sortie du plein écran **avant** d'ouvrir le panneau, des deux côtés : côté iframe (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et** côté parent (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut être détenu par le document parent et non par l'iframe (dans ce cas `document.fullscreenElement` est nul dans l'iframe et sa sortie échoue). Tests : `forge-completion.test.mjs` (+1), `editor-inline.test.mjs` (+1) | Le panneau assistant est monté dans `document.body` du parent : l'API Fullscreen ne rend que l'élément plein écran et ses descendants, donc il ne peut pas s'afficher au-dessus de l'iframe Forge en plein écran. La sortie côté iframe seule ne suffisait pas quand le parent détient le plein écran |
| *BUG-057* | [🟡 IMPORTANT] Assistant IA : le bouton « Ajouter » est inopérant dans l'éditeur Forge (fonctionne seulement dans « Editer ») | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js`, `frontend/editor-poc.html` | Ouvrir un document dans Forge, demander une réponse à l'assistant puis cliquer « Ajouter » | `_insertIntoEditor()` cible Forge (`#forge-iframe`) : `postMessage({ type: 'parent-insert', text })` ; `editor-poc.html` insère au curseur (`insertAtCursor`) et marque le tampon modifié. Repli textarea inclus. Tests : `tests/frontend/ai.test.mjs` (+3), `tests/frontend/editor-inline.test.mjs` (+1) | `state.editorView` (CodeMirror) est nul en Forge : le clic affichait « Aucun document ouvert dans l'éditeur » |
| *BUG-058* | [🔵 MINEUR] Éditeur « Editer » : la barre de numérotation de ligne ne suit pas la couleur du thème (gutter clair `#f5f5f5` en thème sombre) | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css` | Ouvrir un document → Editer en thème sombre : la colonne des numéros de ligne reste gris clair alors que le fond de l'éditeur est sombre | Thème du gutter CodeMirror via les variables CSS (`color-mix(var(--text-primary) …)` pour le fond, `--text-secondary` pour les numéros, `--border` pour la séparation, `--text-primary` pour la ligne active) au lieu des valeurs codées en dur de CodeMirror ; test de non-régression dans `tests/frontend/editor-inline.test.mjs`. Vérifié Playwright (instance de test) : sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de` | CodeMirror applique `background:#f5f5f5` par défaut, indépendamment du thème ObsiGate ; en mode sombre le fond de l'éditeur suit `--bg-secondary` mais pas le gutter |
| *BUG-060* | [🟡 IMPORTANT] Viewer PDF : l'affichage des pages ne fonctionne pas — seule la barre d'outils « PDF — N pages » s'affiche, le contenu reste vide | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/viewer.js`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau) | Cliquer un fichier `.pdf` dans l'arborescence | `frontend/js/viewer.js` : le rendu PDF passe de `<embed type="application/pdf">` à `<iframe>` (autorisée par `frame-src 'self'`, le stream étant same-origin). Tests : `tests/frontend/pdf-viewer.test.mjs` (+6) et `tests/e2e/pdf-viewer.spec.js` (fixture `test_vault/sample-pdf.pdf`) | Cause : la CSP durcie en BUG-034 pose `object-src 'none'`, directive qui gouverne `<embed>`/`<object>` → le lecteur PDF natif était bloqué (barre d'outils rendue, corps vide). Le test E2E échoue bien avec l'ancien `<embed>`. `object-src 'none'` conservé (le correctif ne désarme pas la CSP) |
| | | | | | | | | | | |
### TODOs techniques (améliorations / nouvelles tâches)
@@ -223,6 +227,12 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| 2026-09-17 | #101 | Feature | `frontend/editor-poc.html`, `frontend/js/sync.js`, `frontend/js/viewer.js`, `frontend/js/utils.js`, `frontend/index.html`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/frontend/editor-inline.test.mjs`, `docs/features/forge-assistant.md` (nouvelle), `docs/ROADMAP.md`, `CHANGELOG.md` | **#101** : le bouton « AI Panel » de Forge ouvre désormais l'**Assistant IA** partagé (`postMessage forge-open-ai` → `bookslm.openForCurrentContext()`) au lieu du mini-chat isolé (supprimé) ; Forge lit `localStorage['obsigate_ai_picker']` (`aiPickerSelection()`) pour ses appels `/api/ai/*` et sa complétion fantôme (repli `ollama`), endpoints corrigés (`make-longer`/`make-shorter`, `target_lang`) ; bouton **plein écran** natif ajouté à Forge (iframe `allow="fullscreen"`) et à Editer (`#editor-fullscreen`, conteneur `#editor-container`, sortie à la fermeture, Échap laissé au navigateur) ; i18n `editor.fullscreen`/`editor.exit_fullscreen`. Vérifié : `tests/frontend/editor-inline.test.mjs` 40/40 (+10), unit 9/9, validate-imports 38 modules, 13 suites JSDOM vertes. | 🟢 livré (en attente vérif utilisateur) |
| 2026-09-17 | BUG-055 | Correction | `frontend/editor-poc.html`, `frontend/js/autocomplete.js`, `backend/ai.py`, `.gitea/workflows/ci.yml`, `tests/frontend/forge-completion.test.mjs` (nouveau), `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-055** : trois gestionnaires `keydown` Tab indépendants s'exécutaient à chaque appui — l'indentation (`insertAtCursor(' ')`) s'ajoutait à la complétion de mot (`insertAtCursor(suffixe)`) et à l'acceptation du ghost text, d'où l'espace parasite avant le mot complété puis un effacement destructeur. Gestion **unifiée** de Tab (`liste ouverte > ghost > mot du document > indentation`), helpers purs partagés (`getWordFragment`/`findWordCompletions`/`normalizeGhost`/`chooseTabAction`) extraits dans `autocomplete.js`, liste déroulante au curseur quand plusieurs mots correspondent, ghost **positionné au curseur** (plus de miroir du document entier, nettoyé au déplacement/scroll), complétion de mot sans espace garantie et prompt `/api/ai/inline-complete` simplifié (128 tokens). Vérifié : `tests/frontend/forge-completion.test.mjs` 28/28 (nouveau), `unit.test.mjs` 9/9, `editor-inline.test.mjs` 40/40, `ai.test.mjs` 88/88, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-055 (complément) | Correction | `frontend/editor-poc.html`, `frontend/js/utils.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-055 (complément)** : une complétion acceptée au `Tab` disparaissait 1–2 s plus tard. Cause : l'auto-sauvegarde (2 s) déclenche un `index_updated` SSE sur le fichier affiché, et `reloadExternalWrite` rechargeait le tampon Forge **depuis le disque**, écrasant toute frappe postérieure à la sauvegarde. Le rechargement SSE est désormais ignoré si le tampon est modifié (`parent-reload` sans `force` + `isDirty` ; garde équivalente sur le point d'auto-sauvegarde CodeMirror) ; seul `obsigate:file-written` (assistant IA) passe `force=true`. L'auto-sauvegarde ne remet plus l'état « enregistré » si des modifications sont arrivées pendant la requête (Forge + CodeMirror), et `acceptGhost()` annule la requête de prédiction en attente. Vérifié : `forge-completion.test.mjs` 31/31 (+3), `editor-inline.test.mjs` 41/41 (+1), 14 suites frontend vertes, validate-imports 38 modules. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-056 | Correction | `frontend/editor-poc.html`, `frontend/js/sync.js`, `tests/frontend/forge-completion.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-056** : en plein écran Forge, l'Assistant IA s'ouvrait en arrière-plan. La sortie du plein écran est désormais faite **côté iframe** (`openAssistant` → `document.exitFullscreen()` puis `postMessage` à la résolution) **et côté parent** (`sync.js` sur `forge-open-ai` → `document.exitFullscreen()` puis `openForCurrentContext()`), car le plein écran peut appartenir au document parent (l'iframe voit alors `fullscreenElement` nul et sa sortie échoue — c'était le cas non couvert par le premier correctif). Vérifié : `forge-completion.test.mjs` 32/32 (+1), `editor-inline.test.mjs` 42/42 (+1), 14 suites frontend vertes, validate-imports 38 modules. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-057, #102 | Correction + feature | `frontend/js/bookslm.js`, `frontend/editor-poc.html`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/frontend/ai.test.mjs`, `tests/frontend/editor-inline.test.mjs`, `docs/archive/COMPLETED_v1-v2.md`, `docs/ROADMAP.md`, `CHANGELOG.md` | **BUG-057** : le bouton « Ajouter » de l'assistant ne ciblait que `state.editorView` (CodeMirror) ; en Forge il affichait « Aucun document ouvert dans l'éditeur ». `_insertIntoEditor()` gère désormais les trois surfaces : CodeMirror, l'iframe Forge (`postMessage({ type: 'parent-insert', text })` → `insertAtCursor` dans `editor-poc.html`) et le textarea de repli. **#102** : chaque bloc de code d'une réponse reçoit un bouton « Ajouter la section » (`.bookslm-code-insert`, révélé au survol) qui insère le contenu du bloc sans les délimiteurs ` ``` `. Vérifié : `ai.test.mjs` 91/91 (+3), `editor-inline.test.mjs` 43/43 (+1), `forge-completion.test.mjs` 32/32, unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-058 | Correction | `frontend/style.css`, `tests/frontend/editor-inline.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-058** : la barre de numérotation de ligne de l'éditeur « Editer » ne suivait pas le thème — CodeMirror peint `.cm-gutters` avec des valeurs claires codées en dur (`#f5f5f5`, bordure `#ddd`), visibles en thème sombre. Correctif : le gutter dérive des variables CSS ObsiGate (`background: color-mix(in srgb, var(--text-primary) 5%, transparent)`, `color: var(--text-secondary)`, `border-right: 1px solid var(--border)`, ligne active `color-mix(… 10% …)` / `--text-primary`), donc il suit les 15 thèmes et les 4 modes. Vérifié : `editor-inline.test.mjs` 44/44 (+1), unit 9/9, validate-imports 38 modules, pytest 1101 passed / 6 skipped, ruff 0, mypy 0, et Playwright sur l'instance de test (route `style.css` remplacée par le fichier local) — sombre `color(srgb 0.90 0.93 0.95 / 0.05)` + bordure `#21262d`, clair `color(srgb 0.12 0.14 0.16 / 0.05)` + bordure `#d0d7de`, plus de `rgb(245,245,245)`. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-059 | Correction | `frontend/js/bookslm.js`, `tests/frontend/ai.test.mjs`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-059** : dans une conversation ouverte (post ancré en haut), **tout clic** dans la fenêtre de messages — lien de fichier, étapes, sélection de texte — faisait sauter toute la conversation au bas de la fenêtre. Cause : le gestionnaire `mousedown` de dépintage (prévu pour la molette/tactile/poignée de scroll) se déclenchait aussi sur un simple clic, et le retrait du padding d'ancre (`paddingBottom`) bornait le `scrollTop` à la nouvelle hauteur max → saut au bas. Correctif : helper pur `isScrollbarPress(target, clientX, container)` — un appui ne dépine que s'il vise la **poignée de scroll** (cible = conteneur + zone de gouttière droite) ; molette et tactile conservent leur comportement. Vérifié : `ai.test.mjs` 92/92 (+1), unit 9/9, validate-imports 38 modules, pytest / ruff / mypy inchangés côté backend. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-060 | Correction | `frontend/js/viewer.js`, `.gitea/workflows/ci.yml`, `tests/frontend/pdf-viewer.test.mjs` (nouveau), `tests/e2e/pdf-viewer.spec.js` (nouveau), `test_vault/sample-pdf.pdf` (nouveau), `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **BUG-060** : l'ouverture d'un PDF n'affichait aucune page (barre d'outils « PDF — N pages » présente, corps vide). Cause : la CSP durcie en BUG-034 pose `object-src 'none'` — directive qui gouverne `<embed>`/`<object>` — alors que le viewer rendait le PDF via `<embed type="application/pdf">` : le lecteur natif était bloqué. Correctif : rendu dans une `<iframe>` (autorisée par `frame-src 'self'`, le stream `/api/file/{vault}/pdf/stream` étant same-origin) ; `object-src 'none'` conservé. Tests : `pdf-viewer.test.mjs` 6/6 (statique : pas d'`<embed>`, CSP `frame-src 'self'`, iframe pleine hauteur), `pdf-viewer.spec.js` (E2E : iframe + stream `application/pdf` 200/206 + zéro violation CSP ; échoue bien avec l'ancien `<embed>`). Vérifié : pytest 1184 passed / 6 skipped, frontend 14 suites JSDOM vertes, validate-imports 38 modules, ruff/mypy 0. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-17 | BUG-035, BUG-036, BUG-037, BUG-038, BUG-039, BUG-040 | Correction | `backend/secret_redactor.py`, `backend/collab.py`, `backend/auth/{middleware,password,router}.py`, `backend/indexer.py`, `backend/main.py`, `tests/test_api_main.py`, `tests/test_auth.py`, `tests/test_auth_api.py`, `tests/test_collab.py`, `tests/test_pdf.py`, `CHANGELOG.md`, `docs/ISSUES_TODOLIST.md` | **Lot de 6 bugs mineurs (P2/P3)** : BUG-035 masquage hex conditionné au contexte (git/SHA épargnés) ; BUG-036 jeton WebSocket cookie-only (plus de `?token=`) + plafond de trame 16 Mio ; BUG-037 garde-fou au démarrage (refus d'un bind public sans auth sauf `OBSIGATE_ALLOW_INSECURE=true`) ; BUG-038 Argon2 recalibré 19 Mio/t=2/p=1 ; BUG-039 login uniforme 401 (fini 429/403 distinctifs) ; BUG-040 extraction PDF différée via `enrich_pdf_texts()`. Vérifié : pytest 1204 passed / 6 skipped, ruff 0, mypy 0 (77 fichiers), frontend validate-imports 38 modules + unit 9/9. | 🟢 corrigé (en attente vérif utilisateur) |
---
+8 -22
View File
@@ -1,6 +1,6 @@
# ObsiGate — Roadmap
> **Version :** 2.8.2 | **Dernière mise à jour :** 2026-09-17
> **Version :** 2.11.3 | **Dernière mise à jour :** 2026-09-17
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
> vers les fonctionnalités livrées.
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
@@ -78,23 +78,6 @@
- [x] Personnalisation (clé à molette) : ajouter / supprimer / réordonner les commandes
- [x] i18n FR/EN + tests frontend (helpers purs) + E2E mobile
### 92. Assistant IA — Écosystème d'outils (phase 2 : web étendu, sources connectées, documents)
- **Effort :** 3-5 jours | **Impact :** 🟠 | **Zone :** backend (`backend/tools/`)
- **Dépend de :** #91 (registre + section « steps » + `web_search`/`fetch_url` livrés)
- **Description :** étendre le catalogue d'outils de l'assistant au-delà du vault, en
suivant la feuille de route technique détaillée :
[features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) (frameworks évalués,
bibliothèques par catégorie, transverse retry/cache/secrets/async).
- **Sous-tâches :**
- [x] `web_search` : chaîne de repli sans clé (SearXNG → DuckDuckGo → Bing, `OBSIGATE_WEB_FALLBACK`) — BUG-051
- [ ] `web_search` : fournisseurs optionnels à clé (Tavily, Brave, SerpAPI, Exa)
- [ ] `fetch_url` : pages dynamiques via Playwright (worker isolé) ; crawl multi-pages Scrapy en tâche de fond
- [ ] Sources connectées : Gitea/GitHub (priorité haute) puis Google Drive / OneDrive (OAuth2 `authlib`)
- [ ] Production de documents : conversion, tableurs, PDF/Word (outils WRITE + confirmation)
- [ ] Transverse : `tenacity` (backoff), cache SQLite des résultats web avec TTL, secrets via Infisical
- [ ] Chaque outil : libellé `labels.py` + clés i18n `ai.step.*` FR/EN + tests (httpx mocké)
---
## ⚪ Backlog — Sécurité, architecture & performance (P0/P1)
@@ -199,6 +182,10 @@
| 99 | Sidebar — Filtrage des vues « Récents » et « Sauvegardes » | 2.6.0 | [features/sidebar-filters.md](./features/sidebar-filters.md) |
| 100 | Assistant IA — Deep Research en pastille (au lieu du texte injecté) | 2.6.0 | [features/ai-assistant-history.md](./features/ai-assistant-history.md) |
| 101 | Forge — Assistant IA partagé (bouton AI Panel = assistant, fournisseur/modèle configuré, autocomplétion) + plein écran Forge/Editer | 2.8.0 | [features/forge-assistant.md](./features/forge-assistant.md) |
| BUG-057 | Assistant IA — bouton « Ajouter » fonctionnel dans l'éditeur Forge (en plus d'« Editer ») | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
| 102 | Assistant IA — bouton « Ajouter la section » par bloc de code (insertion du bloc seul) | 2.9.0 | [archive](./archive/COMPLETED_v1-v2.md) |
| 92 | Assistant IA — Écosystème d'outils phase 2 (recherche à clé, cache/retry, Playwright, crawl, Gitea/GitHub, documents XLSX/DOCX/CSV/PDF) | 2.10.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
| 103 | Configuration — clés utilisateur des sources connectées & recherche à clé (page Configurations, `data/api_keys.json`, priorité sur l'env) | 2.11.0 | [features/ai-tools-roadmap.md](./features/ai-tools-roadmap.md) |
---
@@ -206,12 +193,11 @@
| Priorité | Items | Effort total estimé |
|---|---|---|
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100 | ~114 jours réalisés |
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–84, #88–93, #94–100, #102, #92 | ~114 jours réalisés |
| 🔵 P2 restant | #77 Desktop : signature de code (non retenue), 6 tests E2E **manuels** ([protocole](./DESKTOP_E2E_CHECKLIST.md)) | ~0,5-1 jour |
| ⚪ P4 restant | #73 Sync (6-8j) | 6-8 jours |
| ⚪ P2 restant | #92 Assistant IA — écosystème d'outils phase 2 (web étendu, sources connectées, documents) | 3-5 jours |
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (issues BUG-035 → BUG-040) | ~15-23 jours |
| **Total restant** | **10 items + finitions** | **~30-47 jours** |
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (BUG-035 → BUG-040 corrigés) | ~15-23 jours |
| **Total restant** | **7 items + finitions** | **~27-42 jours** |
---
+19
View File
@@ -385,6 +385,25 @@ Fichiers texte non markdown :
---
## #102 — Assistant IA : « Ajouter » dans Forge + ajout d'un bloc de code ✅ TERMINÉ
Deux compléments au bouton « Ajouter » de l'assistant IA.
- **BUG-057 — Forge** : `bookslm.js::_insertIntoEditor()` ne ciblait que
`state.editorView` (CodeMirror de « Editer ») et affichait « Aucun document ouvert
dans l'éditeur » en Forge. Il prend désormais en charge les trois surfaces :
CodeMirror, l'iframe Forge (délégation par `postMessage({ type: 'parent-insert' })`,
insert au curseur via `insertAtCursor` côté `editor-poc.html`) et le textarea de
repli.
- **#102 — Ajout d'un bloc** : chaque bloc de code d'une réponse reçoit un bouton
« Ajouter la section » (révélé au survol, `.bookslm-code-insert`) qui insère
uniquement le contenu du bloc (sans les délimiteurs ` ``` `), au lieu de la réponse
complète.
- **Tests** : `tests/frontend/ai.test.mjs` (+3 : Forge, textarea, bloc de code) ;
`tests/frontend/editor-inline.test.mjs` (+1 : handler `parent-insert`).
---
## Grosses fonctionnalités — fiches dédiées
| # | Feature | Version | Fiche |
+12 -1
View File
@@ -1,6 +1,6 @@
# #92 — Assistant IA — Écosystème d'outils : feuille de route technique
> **Statut :** ⚪ Backlog (phase 1 livrée dans #91)
> **Statut :** ✅ livré (phase 2, version 2.10.0) — phase 1 livrée dans #91
> **Effort estimé :** 3-5 jours pour la phase 2 | **Impact :** 🟠
> **Références :** [Roadmap](../ROADMAP.md) · [Outils & MCP #79](./ai-tools-mcp.md) ·
> [Fenêtre de discussion #91](./ai-assistant-conversation-ux.md) · [Changelog](../../CHANGELOG.md)
@@ -36,6 +36,17 @@ réécriture de la boucle n'est nécessaire.
## 3. Phase 2 — catégories à implémenter
> **Livré (2.10.0, #92).** Récapitulatif des décisions finales :
| Catégorie | Décision livrée |
|---|---|
| Recherche web étendue | Tavily, Brave, SerpAPI, Exa à clé (`OBSIGATE_*_API_KEY`), essayés avant SearXNG ; ordre via `OBSIGATE_WEB_PROVIDERS` |
| Lecture de pages | `fetch_url(render=True)` → worker Playwright isolé (`backend/tools/webrender.py`), dépendance optionnelle + erreur explicite |
| Crawl multi-pages | `crawl_site` (WRITE + confirmation) : BFS httpx borné (≤ 20 pages, même hôte, SSRF sur chaque URL) → condensé Markdown dans le vault. Scrapy écarté (dépendance lourde inutile à cette échelle) |
| Sources connectées | Gitea + GitHub (`git_list_repos`, `git_search_issues`, `git_get_file`) via env/Infisical ; drives cloud (Drive/OneDrive) orientés serveur MCP externe (#79). **#103 (2.11.0)** : les clés (URL Gitea, tokens Gitea/GitHub, clés Tavily/Brave/SerpAPI/Exa) se saisissent aussi dans la page Configurations — `backend/tools/secrets.py`, valeur stockée prioritaire sur l'env |
| Production de documents | `create_xlsx`, `create_docx`, `create_csv`, `create_pdf` — WRITE + confirmation, écrit via `save_raw_file(allow_docs=True)` (path safety + backup) |
| Transverse | Cache SQLite (`webcache.py`, TTL `OBSIGATE_WEB_CACHE_TTL`), retry backoff maison (`OBSIGATE_WEB_RETRY`), secrets par env (Infisical-compatible) |
### 3.1 Recherche web étendue (`web_search`)
- **Fallback sans clé — ✅ livré (BUG-051)** : chaîne de fournisseurs dans
`backend/tools/web.py` — SearXNG auto-hébergé (`OBSIGATE_SEARXNG_URL`) puis, si
+5 -2
View File
@@ -109,11 +109,14 @@ Serveur → client :
## Sécurité
- Authentification obligatoire si `OBSIGATE_AUTH_ENABLED=true` (cookie ou `?token=`).
- Authentification obligatoire si `OBSIGATE_AUTH_ENABLED=true` : le jeton est lu depuis le cookie
HttpOnly `access_token` (envoyé lors du handshake same-origin). Le jeton en query string
(`?token=`) n'est **plus accepté** (BUG-036 : URLs journalisées par les proxies).
- Vérification `check_vault_access()` par connexion (un utilisateur ne peut pas rejoindre une room
d'une vault non autorisée).
- `resolve_safe_path()` empêche toute traversée de chemin (`../../`).
- Bornes anti-abus : `MAX_UPDATE_BYTES` (8 Mo) par mise à jour, `MAX_TEXT_CHARS` (8 Mio) par snapshot.
- Bornes anti-abus : `MAX_UPDATE_BYTES` (8 Mo) par mise à jour, `MAX_TEXT_CHARS` (8 Mio) par snapshot,
`MAX_MESSAGE_CHARS` (16 Mio) par trame brute.
- Le serveur ne décode pas le binaire Yjs : il le stocke et le relaie tel quel (pas de surface
d'attaque supplémentaire côté parsing).
+1
View File
@@ -8,6 +8,7 @@
- **Implémentation réelle (vérifiée 2026-09-07) :**
- **Bugs corrigés (2026-09) :** `api_pdf_stream` crashait en 500 (`NameError: current_user` jamais injecté) ; l'indexation incrémentale du watcher faisait `read_text()` sur les PDFs (garbage) ; Range/206 et `pdf/info` absents malgré le texte ci-dessous.
- **BUG-060 (2026-09-17) :** l'affichage inline ne fonctionnait plus — la CSP durcie en BUG-034 (`object-src 'none'`) bloquait l'`<embed>` du viewer (barre d'outils rendue, corps vide). Le rendu passe par une `<iframe>` (autorisée par `frame-src 'self'`), conforme à E1. Tests : `tests/frontend/pdf-viewer.test.mjs` + `tests/e2e/pdf-viewer.spec.js`.
- `GET /api/file/{vault}/pdf/info` — métadonnées seules sans transférer le document (C3)
- Stream avec `Accept-Ranges` + 206 Partial Content (single range, suffix-range, 416) (C2)
- `OBSIGATE_PDF_MAX_SIZE_MB` (50) + `OBSIGATE_PDF_EXTRACT_TIMEOUT` (30s via thread-pool) (B4/G3)
+24 -3
View File
@@ -1200,9 +1200,23 @@ body { font-family: var(--sans); background: var(--bg); color: var(--text); heig
// iframe, so the AI button simply asks the parent to open it — no duplicated
// panel, and the configured provider/model is shared automatically.
function openAssistant() {
try {
window.parent.postMessage({ type: 'forge-open-ai' }, '*');
} catch (e) { /* not embedded */ }
var notify = function() {
try {
window.parent.postMessage({ type: 'forge-open-ai' }, '*');
} catch (e) { /* not embedded */ }
};
// The shared assistant lives in the parent document, which cannot render
// above a fullscreen Forge iframe: leave fullscreen first so the panel is
// actually visible when it opens.
if (document.fullscreenElement && document.exitFullscreen) {
try {
var p = document.exitFullscreen();
if (p && p.then) p.then(notify, notify);
else notify();
} catch (e) { notify(); }
} else {
notify();
}
}
// ---- Fullscreen (#101) ----
@@ -1460,6 +1474,13 @@ body { font-family: var(--sans); background: var(--bg); color: var(--text); heig
window.addEventListener('message', function(e) {
if (!e.data || !e.data.type) return;
if (e.data.type === 'parent-save') { forceSave(); }
// #102/BUG-057 — the parent AI assistant's « Ajouter » button inserts its
// answer (or a single code block) at the current cursor position.
if (e.data.type === 'parent-insert' && typeof e.data.text === 'string') {
var p = getPos();
insertAtCursor((p.s > 0 ? '\n' : '') + e.data.text);
showToast('Texte ajoute au document', 'ok');
}
// #93 — the parent reloads the document after an external write (AI
// assistant edit_file / append_to_file / create_file): re-read from disk
// and drop the stale local buffer that would otherwise be autosaved back
+134
View File
@@ -1543,6 +1543,7 @@
<li><a href="#cfg-hidden-files" class="help-nav-link" data-i18n="config.section_hidden"></a></li>
<li><a href="#cfg-diags" class="help-nav-link" data-i18n="settings.diagnostics"></a></li>
<li><a href="#cfg-ai" class="help-nav-link" data-i18n="settings.ai"></a></li>
<li><a href="#cfg-sources" class="help-nav-link" data-i18n="config.section_sources">Sources connectées</a></li>
<li><a href="#cfg-themes" class="help-nav-link" data-i18n="settings.themes"></a></li>
<li><a href="#cfg-profile" class="help-nav-link" data-i18n="settings.profile"></a></li>
<li><a href="#cfg-security" class="help-nav-link" data-i18n="settings.security"></a></li>
@@ -2260,6 +2261,133 @@
</div>
</section>
<!-- Connected sources & keyed web search (#103) -->
<section
class="config-section help-section"
id="cfg-sources"
>
<h2 data-i18n="config.section_sources"
>Sources connectées & recherche web</h2
>
<p
class="config-description"
data-i18n="config.sources_desc"
>
Clés utilisées par les outils de l'Assistant IA
(recherche web à clé, Gitea, GitHub). Elles sont
stockées localement et priment sur les variables
d'environnement.
</p>
<div class="config-row">
<label
class="config-label"
for="cfg-tavily-key"
>Tavily API Key</label
>
<input
type="password"
id="cfg-tavily-key"
class="config-input"
placeholder="tvly-..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-brave-key"
>Brave Search API Key</label
>
<input
type="password"
id="cfg-brave-key"
class="config-input"
placeholder="BSA..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-serpapi-key"
>SerpAPI Key</label
>
<input
type="password"
id="cfg-serpapi-key"
class="config-input"
placeholder="..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-exa-key"
>Exa API Key</label
>
<input
type="password"
id="cfg-exa-key"
class="config-input"
placeholder="..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-gitea-url"
data-i18n="config.gitea_url"
>URL Gitea</label
>
<input
type="text"
id="cfg-gitea-url"
class="config-input"
placeholder="https://git.example.net"
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-gitea-token"
>Gitea Token</label
>
<input
type="password"
id="cfg-gitea-token"
class="config-input"
placeholder="token personnel..."
autocomplete="off"
/>
</div>
<div class="config-row">
<label
class="config-label"
for="cfg-github-token"
>GitHub Token</label
>
<input
type="password"
id="cfg-github-token"
class="config-input"
placeholder="ghp_..."
autocomplete="off"
/>
</div>
<div
class="config-actions-row"
style="margin-top: 16px"
>
<button class="config-btn-save"
id="cfg-save-tool-keys" data-i18n="help.shortcut_save">
Sauvegarder
</button>
</div>
</section>
<!-- Themes -->
<section
class="config-section help-section"
@@ -4319,6 +4447,12 @@
Réponses formatées : titres, listes, tableaux, citations et
blocs de code.
</li>
<li data-i18n="help.assistant_insert">
Le bouton « Ajouter » (au survol d'une réponse) insère la
réponse dans le document ouvert dans l'éditeur (Editer ou
Forge) ; chaque bloc de code propose « Ajouter la section »
pour n'insérer que ce bloc.
</li>
<li data-i18n="help.assistant_links">
Les fichiers et chemins cités sont des liens : cliquez sur un
fichier pour l'ouvrir, sur un dossier pour le révéler dans
+104 -17
View File
@@ -43,6 +43,25 @@ const PANEL_MIN_WIDTH = 320;
const PANEL_MAX_WIDTH = 1000;
const PANEL_WIDTH_KEY = 'obsigate-bookslm-width';
/**
* BUG-059 — Is this pointer press a scrollbar drag (and only that)?
*
* Only genuine scroll gestures may release the top-pinning of the latest
* question (wheel / touchmove / scrollbar drag). A scrollbar drag reports the
* scrollable container itself as the event target and lands inside the
* vertical scrollbar gutter. A plain click anywhere in the content — a link,
* a file path, the steps toggle, text selection — must NOT unpin: clearing
* the anchor padding would clamp the scroll position and jump the whole
* thread to the bottom of the conversation window.
*/
export function isScrollbarPress(target, clientX, container) {
if (!container || target !== container) return false;
const rect = container.getBoundingClientRect();
if (!rect || !Number.isFinite(rect.right)) return false;
const GUTTER = 24; // conservative vertical-scrollbar width estimate
return clientX >= rect.right - GUTTER;
}
/**
* Accent- and case-insensitive normalization used to match `/` commands and
* `@` mentions: skill ids/labels and vault paths may contain accented
@@ -985,7 +1004,12 @@ class BooksLM {
};
messagesEl.addEventListener('wheel', unpin, { passive: true });
messagesEl.addEventListener('touchmove', unpin, { passive: true });
messagesEl.addEventListener('mousedown', unpin, { passive: true });
// BUG-059: a scrollbar drag only — a plain click on the content must
// not unpin (it would clear the anchor padding and jump the thread to
// the bottom of the window).
messagesEl.addEventListener('mousedown', (e) => {
if (isScrollbarPress(e.target, e.clientX, messagesEl)) unpin();
}, { passive: true });
}
// Close the session / command menus when clicking elsewhere in the panel.
@@ -2153,6 +2177,7 @@ class BooksLM {
bubble.className = 'bookslm-bubble assistant';
const { text, actions } = this._extractActions(msg.content || '');
bubble.innerHTML = this._renderMarkdown(text);
this._enhanceCodeBlocks(bubble);
// Plain chat has no steps block: the pending answer itself carries the
// animated indicator until the first token arrives.
if (running && !text && !actions.length) {
@@ -2234,28 +2259,90 @@ class BooksLM {
}
/**
* Append an assistant answer to the Forge editor document (Notion “Ajouter”).
* No-ops with a toast when no editor session is open.
* Append an assistant answer (or a single code block) to the document open
* in the active editor (Notion “Ajouter”). Supports the three editor
* surfaces: CodeMirror (« Editer »), the Forge iframe and the plain textarea
* fallback. No-ops with a toast when no editor session is open.
*/
_insertIntoEditor(text) {
const value = String(text == null ? '' : text);
if (!value.trim()) return;
// 1. CodeMirror editor (« Editer »).
const view = state.editorView;
if (!view || !view.state || typeof view.dispatch !== 'function') {
showToast(t('bookslm.insert_no_editor'), 'info');
if (view && view.state && typeof view.dispatch === 'function') {
try {
const at = view.state.selection.main.to;
const insert = (at > 0 ? '\n' : '') + value;
view.dispatch({
changes: { from: at, insert },
selection: { anchor: at + insert.length },
});
if (typeof view.focus === 'function') view.focus();
showToast(t('bookslm.inserted'), 'success');
} catch (e) {
console.warn('AI assistant: insert into editor failed', e);
showToast(t('bookslm.insert_no_editor'), 'error');
}
return;
}
try {
const at = view.state.selection.main.to;
const insert = (at > 0 ? '\n' : '') + text;
view.dispatch({
changes: { from: at, insert },
selection: { anchor: at + insert.length },
});
if (typeof view.focus === 'function') view.focus();
showToast(t('bookslm.inserted'), 'success');
} catch (e) {
console.warn('AI assistant: insert into editor failed', e);
showToast(t('bookslm.insert_no_editor'), 'error');
// 2. Forge editor (same-origin iframe): its buffer lives in the child
// document, so the insertion is delegated with a postMessage.
const forge = document.getElementById('forge-iframe');
if (forge && forge.contentWindow) {
try {
forge.contentWindow.postMessage({ type: 'parent-insert', text: value }, '*');
showToast(t('bookslm.inserted'), 'success');
} catch (e) {
console.warn('AI assistant: insert into Forge failed', e);
showToast(t('bookslm.insert_no_editor'), 'error');
}
return;
}
// 3. Plain textarea fallback (CodeMirror failed to load).
const ta = state.fallbackEditorEl;
if (ta && typeof ta.value === 'string') {
const start = ta.selectionStart != null ? ta.selectionStart : ta.value.length;
const end = ta.selectionEnd != null ? ta.selectionEnd : ta.value.length;
const insert = (start > 0 ? '\n' : '') + value;
ta.value = ta.value.slice(0, start) + insert + ta.value.slice(end);
const pos = start + insert.length;
if (typeof ta.setSelectionRange === 'function') ta.setSelectionRange(pos, pos);
ta.dispatchEvent(new Event('input', { bubbles: true }));
if (typeof ta.focus === 'function') ta.focus();
showToast(t('bookslm.inserted'), 'success');
return;
}
showToast(t('bookslm.insert_no_editor'), 'info');
}
/**
* #102 — Add a discreet « Ajouter » button to every fenced code block of an
* answer, so a single proposed section can be inserted instead of the whole
* reply. The button reuses the message action styling.
*/
_enhanceCodeBlocks(bubble) {
if (!bubble || typeof bubble.querySelectorAll !== 'function') return;
const blocks = bubble.querySelectorAll('pre');
blocks.forEach((pre) => {
const code = pre.querySelector('code');
if (!code) return;
const wrap = document.createElement('div');
wrap.className = 'bookslm-code-block';
pre.parentNode.insertBefore(wrap, pre);
wrap.appendChild(pre);
const btn = this._actionBtn(
t('bookslm.insert_block'),
'insert',
t('bookslm.insert_block_hint'),
() => this._insertIntoEditor(code.textContent),
);
btn.classList.add('bookslm-code-insert');
wrap.appendChild(btn);
});
}
_copyText(text) {
+100
View File
@@ -715,6 +715,7 @@ function initConfigModal() {
await loadHiddenFilesSettings();
loadWebhooksUI();
loadSharesUI();
loadToolKeys();
safeCreateIcons();
});
@@ -765,6 +766,9 @@ function initConfigModal() {
if (saveAIKeysBtn) saveAIKeysBtn.addEventListener("click", saveAIKeys);
const testAIKeysBtn = document.getElementById("cfg-test-ai-keys");
if (testAIKeysBtn) testAIKeysBtn.addEventListener("click", testAIKeys);
// Tool & connected-source keys (#103)
const saveToolKeysBtn = document.getElementById("cfg-save-tool-keys");
if (saveToolKeysBtn) saveToolKeysBtn.addEventListener("click", saveToolKeys);
// Default provider/model selection
const aiDefaultProviderSel = document.getElementById("cfg-ai-default-provider");
if (aiDefaultProviderSel) {
@@ -1715,6 +1719,102 @@ async function testAIKeys() {
}
// ── Tool & connected-source keys (#103) ──
const TOOL_KEY_MAP = {
"cfg-tavily-key": "OBSIGATE_TAVILY_API_KEY",
"cfg-brave-key": "OBSIGATE_BRAVE_API_KEY",
"cfg-serpapi-key": "OBSIGATE_SERPAPI_API_KEY",
"cfg-exa-key": "OBSIGATE_EXA_API_KEY",
"cfg-gitea-url": "OBSIGATE_GITEA_URL",
"cfg-gitea-token": "OBSIGATE_GITEA_TOKEN",
"cfg-github-token": "OBSIGATE_GITHUB_TOKEN",
};
function _ensureToolKeyUI() {
for (const [inputId] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(inputId);
if (!input) continue;
const row = input.closest(".config-row");
if (!row || row.dataset.toolKeyEnhanced) continue;
row.dataset.toolKeyEnhanced = "1";
row.style.cssText += "display:flex;align-items:center;gap:8px;flex-wrap:wrap;";
const badge = document.createElement("span");
badge.id = inputId + "-badge";
badge.style.cssText = "font-size:11px;padding:2px 8px;border-radius:10px;white-space:nowrap;";
row.appendChild(badge);
const delBtn = document.createElement("button");
delBtn.type = "button";
delBtn.id = inputId + "-delete";
delBtn.className = "config-btn-secondary";
delBtn.style.cssText = "font-size:11px;padding:4px 10px;color:var(--danger,#e74c3c);border-color:var(--danger,#e74c3c);cursor:pointer;display:none;";
delBtn.textContent = "\u00d7 " + t("config.delete_key");
delBtn.addEventListener("click", () => deleteToolKey(inputId));
row.appendChild(delBtn);
}
}
function _setToolKeyBadge(inputId, hasKey) {
const badge = document.getElementById(inputId + "-badge");
const delBtn = document.getElementById(inputId + "-delete");
if (badge) {
if (hasKey) {
badge.textContent = "\u2713 " + t("config.key_set");
badge.style.background = "var(--success-bg, #27ae6022)";
badge.style.color = "var(--success, #27ae60)";
badge.style.border = "1px solid var(--success, #27ae60)";
} else {
badge.textContent = t("config.key_unset");
badge.style.background = "var(--muted-bg, #ffffff10)";
badge.style.color = "var(--text-muted, #888)";
badge.style.border = "1px solid var(--border, #444)";
}
}
if (delBtn) delBtn.style.display = hasKey ? "inline-block" : "none";
}
async function loadToolKeys() {
_ensureToolKeyUI();
try {
const data = await api("/api/config/tool-keys");
for (const [inputId, name] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(inputId);
const val = data[name] || "";
if (input && !input.value.trim()) input.placeholder = val || input.placeholder;
_setToolKeyBadge(inputId, !!val);
}
} catch(e) { /* non-admin: the section stays inert */ }
}
async function saveToolKeys() {
const keys = {};
for (const [id, name] of Object.entries(TOOL_KEY_MAP)) {
const input = document.getElementById(id);
if (input && input.value.trim()) keys[name] = input.value.trim();
}
if (!Object.keys(keys).length) {
showToast(t("config.no_keys"), "warning");
return;
}
try {
await api("/api/config/tool-keys", { method: "POST", body: JSON.stringify(keys) });
showToast(t("config.api_keys_saved"), "success");
Object.keys(TOOL_KEY_MAP).forEach(id => { const el = document.getElementById(id); if (el) el.value = ""; });
loadToolKeys();
} catch(e) { showToast("Erreur: " + e.message, "error"); }
}
async function deleteToolKey(inputId) {
const name = TOOL_KEY_MAP[inputId];
if (!name) return;
if (!confirm(t("config.delete_key_confirm") + " " + name + " ?")) return;
try {
await api("/api/config/tool-keys/" + name, { method: "DELETE" });
showToast(t("config.key_deleted") + " " + name, "success");
loadToolKeys();
} catch(e) { showToast("Erreur: " + e.message, "error"); }
}
export {
initSidebarTabs,
initConfigModal,
+20 -5
View File
@@ -649,11 +649,26 @@ export function init() {
// #101 — Forge has no rich AI panel of its own: its AI button asks the
// parent to open the shared AI Assistant (same provider/model, skills,
// context and history).
import('./bookslm.js').then(function(m) {
if (m && m.default) m.default.openForCurrentContext();
}).catch(function(err) {
console.warn('[Forge] Unable to open AI assistant', err);
});
var openForgeAssistant = function() {
import('./bookslm.js').then(function(m) {
if (m && m.default) m.default.openForCurrentContext();
}).catch(function(err) {
console.warn('[Forge] Unable to open AI assistant', err);
});
};
// BUG-056 — the assistant panel is mounted in this document. The native
// fullscreen may be owned by the parent (Forge iframe) rather than by the
// iframe itself, so the parent must also leave fullscreen before the
// panel can be seen.
if (document.fullscreenElement && document.exitFullscreen) {
try {
document.exitFullscreen().catch(function() {}).then(openForgeAssistant, openForgeAssistant);
} catch (err) {
openForgeAssistant();
}
} else {
openForgeAssistant();
}
}
});
}
+1 -1
View File
@@ -555,7 +555,7 @@ export function renderFile(data) {
</div>
<div class="pdf-body">
${tocHtml}
<embed src="${pdfUrl}" class="pdf-iframe" type="application/pdf" title="${escapeHtml(data.title)}"></embed>
<iframe src="${pdfUrl}" class="pdf-iframe" title="${escapeHtml(data.title)}"></iframe>
</div>
</div>`;
lucide.createIcons();
+19
View File
@@ -364,6 +364,14 @@
"config.ai_model": "Model",
"config.ai_openrouter_label": "OpenRouter API Key",
"config.api_keys_saved": "API keys saved",
"config.section_sources": "🔗 Connected sources & search",
"config.sources_desc": "Keys used by the AI assistant tools (keyed web search: Tavily, Brave, SerpAPI, Exa; connected sources: Gitea, GitHub). They are stored server-side and take precedence over environment variables.",
"config.gitea_url": "Gitea URL",
"config.key_set": "Configured",
"config.key_unset": "Not configured",
"config.delete_key": "Delete",
"config.delete_key_confirm": "Delete key",
"config.key_deleted": "Key deleted:",
"config.backups": "Backups",
"config.backups_desc": "Manage automatic file backups.",
"config.client_config": "Client config",
@@ -1290,6 +1298,7 @@
"help.assistant_panel": "🧠 Assistant panel (BooksLM)",
"help.assistant_panel_desc": "The side assistant (floating button or a folder's context menu) answers in formatted Markdown and contextualises your directories or documents. In general mode it also knows what you are looking at: open documents, current directory, active search and recently modified files.",
"help.assistant_markdown": "Formatted answers: headings, lists, tables, quotes and code blocks.",
"help.assistant_insert": "The \"Add\" button (revealed on hover of an answer) inserts the answer into the document open in the editor (Editer or Forge); each code block offers \"Add section\" to insert just that block.",
"help.assistant_links": "Cited files and paths are links: a bare filename copies the name to the clipboard, a folder is revealed in the tree, and a file path opens it in the viewer.",
"help.assistant_sessions": "The header history icon lists past sessions (reopen or delete); “+” starts a new conversation.",
"help.assistant_agent": "The \"agent mode\" button enables tools (read, list, search); modifying actions require confirmation with a change preview.",
@@ -1785,6 +1794,8 @@
"bookslm.insert_hint": "Append the answer to the document open in the editor",
"bookslm.inserted": "Answer added to the document",
"bookslm.insert_no_editor": "No document open in the editor",
"bookslm.insert_block": "Add section",
"bookslm.insert_block_hint": "Add only this code block to the document open in the editor",
"ai.steps_count": "{count} step",
"ai.steps_count_plural": "{count} steps",
"ai.activity_thinking": "Thinking…",
@@ -1819,6 +1830,14 @@
"ai.step.vaults": "Listed the vaults",
"ai.step.fetch_url": "Opened a web page: {value}",
"ai.step.web_search": "Searched the web: {value}",
"ai.step.crawl": "Crawled a site: {value}",
"ai.step.git_repos": "Listed repositories ({value})",
"ai.step.git_issues": "Searched issues: {value}",
"ai.step.git_file": "Read a repo file: {value}",
"ai.step.xlsx_create": "Spreadsheet proposed: {value}",
"ai.step.docx_create": "Word document proposed: {value}",
"ai.step.csv_create": "CSV file proposed: {value}",
"ai.step.pdf_create": "PDF document proposed: {value}",
"bookslm.copied": "Copied to clipboard",
"bookslm.error": "AI service error",
"bookslm.regenerate": "Regenerate",
+19
View File
@@ -364,6 +364,14 @@
"config.ai_model": "Modèle",
"config.ai_openrouter_label": "OpenRouter API Key",
"config.api_keys_saved": "Clés API sauvegardées",
"config.section_sources": "🔗 Sources connectées & recherche",
"config.sources_desc": "Clés utilisées par les outils de l'Assistant IA (recherche web à clé : Tavily, Brave, SerpAPI, Exa ; sources connectées : Gitea, GitHub). Elles sont stockées sur le serveur et priment sur les variables d'environnement.",
"config.gitea_url": "URL Gitea",
"config.key_set": "Configuré",
"config.key_unset": "Non configuré",
"config.delete_key": "Supprimer",
"config.delete_key_confirm": "Supprimer la clé",
"config.key_deleted": "Clé supprimée :",
"config.backups": "Sauvegardes",
"config.backups_desc": "Gérez les sauvegardes automatiques de vos fichiers.",
"config.client_config": "Configuration client",
@@ -1290,6 +1298,7 @@
"help.assistant_panel": "🧠 Panneau Assistant (BooksLM)",
"help.assistant_panel_desc": "L'assistant latéral (bouton flottant ou menu contextuel d'un dossier) répond en Markdown formaté et contextualise vos répertoires ou documents. En contexte général, il connaît aussi ce que vous voyez : documents ouverts, répertoire courant, recherche en cours et fichiers récemment modifiés.",
"help.assistant_markdown": "Réponses formatées : titres, listes, tableaux, citations et blocs de code.",
"help.assistant_insert": "Le bouton « Ajouter » (au survol d'une réponse) insère la réponse dans le document ouvert dans l'éditeur (Editer ou Forge) ; chaque bloc de code propose « Ajouter la section » pour n'insérer que ce bloc.",
"help.assistant_links": "Les fichiers et chemins cités sont des liens : un simple nom de fichier copie le nom dans le presse-papiers, un dossier est révélé dans l'arborescence, et un chemin de fichier l'ouvre dans le viewer.",
"help.assistant_sessions": "L'icône historique de l'en-tête liste les sessions passées (recharger ou supprimer) ; « + » démarre une nouvelle conversation.",
"help.assistant_agent": "Le bouton « mode agent » active les outils (lire, lister, chercher) ; les actions de modification demandent une confirmation avec aperçu des changements.",
@@ -1785,6 +1794,8 @@
"bookslm.insert_hint": "Ajouter la réponse au document ouvert dans l'éditeur",
"bookslm.inserted": "Réponse ajoutée au document",
"bookslm.insert_no_editor": "Aucun document ouvert dans l'éditeur",
"bookslm.insert_block": "Ajouter la section",
"bookslm.insert_block_hint": "Ajouter uniquement ce bloc de code au document ouvert dans l'éditeur",
"ai.steps_count": "{count} étape",
"ai.steps_count_plural": "{count} étapes",
"ai.activity_thinking": "Réflexion…",
@@ -1819,6 +1830,14 @@
"ai.step.vaults": "Liste des vaults consultée",
"ai.step.fetch_url": "Page web consultée : {value}",
"ai.step.web_search": "Recherche sur le web : {value}",
"ai.step.crawl": "Site exploré : {value}",
"ai.step.git_repos": "Dépôts listés ({value})",
"ai.step.git_issues": "Issues recherchées : {value}",
"ai.step.git_file": "Fichier de dépôt lu : {value}",
"ai.step.xlsx_create": "Tableur proposé : {value}",
"ai.step.docx_create": "Document Word proposé : {value}",
"ai.step.csv_create": "Fichier CSV proposé : {value}",
"ai.step.pdf_create": "Document PDF proposé : {value}",
"bookslm.copied": "Réponse copiée dans le presse-papiers",
"bookslm.error": "Erreur du service AI",
"bookslm.regenerate": "Régénérer",
+22
View File
@@ -3592,6 +3592,20 @@ select {
overflow-x: auto !important;
max-width: 100%;
}
/* BUG-058: CodeMirror paints its line-number gutter with hardcoded light
defaults (#f5f5f5 background, #ddd border), so the bar stayed pale in dark
themes while the editor body followed the theme. Deriving the tint from
`--text-primary` keeps the gutter subtle and makes it adapt to every theme
and mode (dark, light, high-contrast, sepia). */
.cm-editor .cm-gutters {
background: color-mix(in srgb, var(--text-primary) 5%, transparent) !important;
color: var(--text-secondary) !important;
border-right: 1px solid var(--border) !important;
}
.cm-editor .cm-gutters .cm-activeLineGutter {
background: color-mix(in srgb, var(--text-primary) 10%, transparent) !important;
color: var(--text-primary) !important;
}
.fallback-editor {
width: 100%;
min-height: 100%;
@@ -9531,6 +9545,14 @@ body.popup-mode .content-area {
.bookslm-bubble.assistant { align-self: flex-start; background: transparent; padding: 0; width: 100%; max-width: 100%; color: var(--text-primary); }
.bookslm-bubble.assistant code { background: rgba(0,0,0,0.2); padding: 1px 4px; border-radius: 3px; font-size: 0.9em; }
.bookslm-bubble.assistant pre { background: rgba(0,0,0,0.3); padding: 10px; border-radius: 6px; overflow-x: auto; margin: 8px 0; }
/* #102 — per-code-block “Ajouter”: a discreet button in the block's corner,
revealed on hover/focus, that inserts only this section. */
.bookslm-code-block { position: relative; }
.bookslm-code-block .bookslm-code-insert { position: absolute; top: 6px; right: 6px;
z-index: 1; background: var(--surface2); border-color: var(--border); opacity: 0;
transition: opacity 0.15s; }
.bookslm-code-block:hover .bookslm-code-insert,
.bookslm-code-block:focus-within .bookslm-code-insert { opacity: 1; }
.bookslm-sources { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 8px; }
.bookslm-source-badge { display: inline-flex; align-items: center; gap: 4px; padding: 2px 8px;
background: var(--surface); border: 1px solid var(--border); border-radius: 12px;
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "obsigate",
"version": "2.8.2",
"version": "2.11.3",
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
"main": "patch.js",
"directories": {
+106
View File
@@ -0,0 +1,106 @@
%PDF-1.3
%“Œ‹ž ReportLab Generated PDF document (opensource)
1 0 obj
<<
/F1 2 0 R
>>
endobj
2 0 obj
<<
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
>>
endobj
3 0 obj
<<
/Contents 9 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
4 0 obj
<<
/Contents 10 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
5 0 obj
<<
/Contents 11 0 R /MediaBox [ 0 0 612 792 ] /Parent 8 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
6 0 obj
<<
/PageMode /UseNone /Pages 8 0 R /Type /Catalog
>>
endobj
7 0 obj
<<
/Author (anonymous) /CreationDate (D:20260917153218-04'00') /Creator (anonymous) /Keywords () /ModDate (D:20260917153218-04'00') /Producer (ReportLab PDF Library - \(opensource\))
/Subject (unspecified) /Title (untitled) /Trapped /False
>>
endobj
8 0 obj
<<
/Count 3 /Kids [ 3 0 R 4 0 R 5 0 R ] /Type /Pages
>>
endobj
9 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jt]&.8<uSu(.bn9(BF;Q*M*~>endstream
endobj
10 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?Jruos8<uSu(.bn9(BF;_*M3~>endstream
endobj
11 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 135
>>
stream
GapQh0E=F,0U\H3T\pNYT^QKk?tc>IP,;W#U1^23ihPEM_?CU^!/VX4lBrg6%A?7Y)Zc(P6:e82L4<*@VL)cDW^;5oRmL:j77h2jWe.B?6"5/?K!D1>8<uSu(.bn9(BF;m*M<~>endstream
endobj
xref
0 12
0000000000 65535 f
0000000061 00000 n
0000000092 00000 n
0000000199 00000 n
0000000392 00000 n
0000000586 00000 n
0000000780 00000 n
0000000848 00000 n
0000001109 00000 n
0000001180 00000 n
0000001405 00000 n
0000001631 00000 n
trailer
<<
/ID
[<464fc7cfdf793a5b6d29e3d0d043c5a7><464fc7cfdf793a5b6d29e3d0d043c5a7>]
% ReportLab generated PDF document -- digest (opensource)
/Info 7 0 R
/Root 6 0 R
/Size 12
>>
startxref
1857
%%EOF
+17
View File
@@ -22,6 +22,23 @@ def _reset_tool_ratelimit():
ratelimit.reset()
@pytest.fixture(autouse=True)
def _disable_web_cache():
"""Web cache off by default: tests stay hermetic (no cross-test hits).
tests/test_web_cache.py re-enables it explicitly with a tmp path.
"""
from backend.tools import webcache
saved_path = os.environ.get("OBSIGATE_WEB_CACHE_PATH")
os.environ["OBSIGATE_WEB_CACHE_TTL"] = "0"
yield
if saved_path is None:
os.environ.pop("OBSIGATE_WEB_CACHE_PATH", None)
else:
os.environ["OBSIGATE_WEB_CACHE_PATH"] = saved_path
@pytest.fixture(autouse=True)
def _clean_env():
"""Ensure no vault env vars leak between tests — but preserve test vault config."""
+92
View File
@@ -0,0 +1,92 @@
/**
* E2E tests for ObsiGate PDF inline viewer (BUG-060).
*
* Regression : la CSP posée par BUG-034 (`object-src 'none'`) bloque l'élément
* `<embed>` qui servait le PDF. Le viewer doit rendre le stream dans une
* `<iframe>` (autorisée par `frame-src 'self'`).
*
* Fixtures : `test_vault/sample-pdf.pdf` (3 pages, texte simple).
*
* Run (local):
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js
* BASE_URL=http://localhost:2029 npx playwright test tests/e2e/pdf-viewer.spec.js --headed
*/
import { test, expect } from '@playwright/test';
const BASE = process.env.BASE_URL || 'http://localhost:2029';
const CREDS = {
username: process.env.OBSIGATE_USER || 'admin',
password: process.env.OBSIGATE_PASS || 'test123',
};
async function login(page) {
await page.goto(BASE);
const loginForm = page.locator('#login-screen');
await expect(loginForm).toBeVisible({ timeout: 5000 }).catch(() => {});
if (await loginForm.isVisible()) {
await page.fill('#login-username', CREDS.username);
await page.fill('#login-password', CREDS.password);
await page.click('#login-btn');
}
await page.waitForFunction(() => window.__OBSIGATE_BOOTED === true, { timeout: 20000 });
}
async function openFile(page, vault, filePath) {
const treeItem = page.locator(`.tree-item[data-vault="${vault}"][data-path="${filePath}"]`);
// Le vault peut être replié (auth activée) : l'étendre avant de chercher le fichier.
if (!(await treeItem.count())) {
await page.locator(`.tree-item.vault-item[data-vault="${vault}"]`).first().click();
await treeItem.waitFor({ state: 'attached', timeout: 8000 });
}
await treeItem.dblclick({ timeout: 5000 });
await page.waitForTimeout(500);
}
test.describe('PDF viewer — affichage inline (BUG-060)', () => {
test('ouvre un PDF dans une iframe (pas d\'<embed>) et le stream charge sans violation CSP', async ({ page }) => {
const cspViolations = [];
page.on('console', (msg) => {
const text = msg.text();
if (text.includes('Content Security Policy') && (text.includes('object-src') || text.includes('Refused'))) {
cspViolations.push(text);
}
});
await login(page);
// Le stream est demandé au moment où le viewer monte l'iframe : enregistrer
// l'écoute AVANT d'ouvrir le fichier (sinon la réponse est déjà passée).
const streamResponsePromise = page.waitForResponse(
(r) => r.url().includes('/pdf/stream') && (r.status() === 200 || r.status() === 206),
{ timeout: 15000 },
);
await openFile(page, 'TestVault', 'sample-pdf.pdf');
const iframe = page.locator('#content-area .pdf-iframe');
await expect(iframe).toBeVisible({ timeout: 15000 });
await expect(iframe).toHaveAttribute('src', /\/api\/file\/TestVault\/pdf\/stream\?path=/);
// La balise doit être une iframe (le <embed>/<object> serait bloqué par CSP)
const tagName = await iframe.evaluate((el) => el.tagName);
expect(tagName).toBe('IFRAME');
expect(await page.locator('#content-area embed, #content-area object').count()).toBe(0);
// La barre d'outils indique le nombre de pages du PDF
await expect(page.locator('.pdf-info')).toContainText('3 pages');
// Le stream est bien servi en application/pdf (200 ou 206 Range)
const streamResp = await streamResponsePromise;
expect(streamResp.headers()['content-type'] || '').toContain('application/pdf');
// Le cadre embarque réellement le document PDF (navigateur natif)
const pdfFrame = await iframe.contentFrame();
expect(pdfFrame).not.toBeNull();
// Aucune violation CSP liée à object-src pendant l'ouverture
expect(cspViolations).toEqual([]);
});
});
+86
View File
@@ -302,6 +302,25 @@ async function main() {
}
});
await test("isScrollbarPress: content clicks never unpin, scrollbar drags do (BUG-059)", () => {
const { isScrollbarPress } = bookslmMod;
const container = document.createElement("div");
container.getBoundingClientRect = () => ({
left: 0, right: 800, top: 0, bottom: 600, width: 800, height: 600,
});
const link = document.createElement("a");
// Click on a file link / steps toggle → must NOT unpin (the thread
// would otherwise jump to the bottom when the padding is cleared).
assert.equal(isScrollbarPress(link, 790, container), false);
// Click on blank content area → must NOT unpin either.
assert.equal(isScrollbarPress(container, 300, container), false);
// Dragging the vertical scrollbar: target is the container itself and
// the press lands in the right-edge gutter → unpin allowed.
assert.equal(isScrollbarPress(container, 792, container), true);
// Defensive: no container / no geometry → never unpin.
assert.equal(isScrollbarPress(link, 790, null), false);
});
await test("the placeholder render keeps the anchor (no scroll reset before first token)", async () => {
// Regression: the assistant placeholder used to be rendered while
// `_isLoading` was still false, so it took the "restore previous
@@ -404,6 +423,73 @@ async function main() {
panel.remove();
});
await test("add action delegates to the Forge iframe when no CodeMirror is open (BUG-057)", () => {
const b = new BooksLM();
const panel = b._render();
b._panel = panel;
document.body.appendChild(panel);
b._messages = [{ role: "assistant", content: "Réponse Forge" }];
b._renderMessages();
state.editorView = null;
const iframe = document.createElement("iframe");
iframe.id = "forge-iframe";
document.body.appendChild(iframe);
const posted = [];
iframe.contentWindow.postMessage = (msg) => posted.push(msg);
panel.querySelector(".bookslm-msg.assistant .bookslm-msg-action-insert").click();
assert.equal(posted.length, 1, "one postMessage to the Forge iframe");
assert.equal(posted[0].type, "parent-insert");
assert.ok(posted[0].text.includes("Réponse Forge"), "answer forwarded");
iframe.remove();
panel.remove();
});
await test("add action falls back to the plain textarea editor", () => {
const b = new BooksLM();
const panel = b._render();
b._panel = panel;
document.body.appendChild(panel);
b._messages = [{ role: "assistant", content: "Réponse textarea" }];
b._renderMessages();
state.editorView = null;
const ta = document.createElement("textarea");
ta.value = "Ligne existante";
document.body.appendChild(ta);
ta.setSelectionRange(ta.value.length, ta.value.length);
state.fallbackEditorEl = ta;
panel.querySelector(".bookslm-msg.assistant .bookslm-msg-action-insert").click();
assert.ok(ta.value.includes("Réponse textarea"), "answer appended to the textarea");
state.fallbackEditorEl = null;
ta.remove();
panel.remove();
});
await test("code block exposes an add-section button inserting only the block (#102)", () => {
const b = new BooksLM();
const panel = b._render();
b._panel = panel;
document.body.appendChild(panel);
b._messages = [{
role: "assistant",
content: "Voici la section :\n\n```markdown\n## Titre\n\nContenu\n```\n\nFin.",
}];
b._renderMessages();
const btn = panel.querySelector(".bookslm-msg.assistant .bookslm-code-insert");
assert.ok(btn, "per-block add button rendered");
const dispatched = [];
state.editorView = {
state: { selection: { main: { to: 0 } } },
dispatch: (spec) => dispatched.push(spec),
focus: () => {},
};
btn.click();
assert.equal(dispatched.length, 1, "editor dispatch called");
assert.ok(dispatched[0].changes.insert.includes("## Titre"), "block content inserted");
assert.ok(!dispatched[0].changes.insert.includes("```"), "code fences stripped");
state.editorView = null;
panel.remove();
});
await test("manual wheel scroll releases the top-pinning", async () => {
const b = new BooksLM();
const panel = b._render();
+28
View File
@@ -299,6 +299,14 @@ test("sync.js forge-close goes through the shared close path", () => {
assert.match(handler[1], /closeEditor\(\);/);
});
test("sync.js leaves fullscreen before opening the Forge assistant (BUG-056)", () => {
const handler = syncSrc.match(/if \(e\.data\.type === 'forge-open-ai'\) \{([\s\S]*?)\n \}/);
assert.ok(handler, "forge-open-ai handler not found");
assert.match(handler[1], /document\.fullscreenElement/);
assert.match(handler[1], /document\.exitFullscreen\(\)/);
assert.match(handler[1], /openForCurrentContext\(\)/);
});
test("sync.js keeps an open edition session alive on external file changes", () => {
const sse = syncSrc.match(/const changed = \(data\.changes \|\| \[\]\)([\s\S]*?)\n \}/);
assert.ok(sse, "SSE index_updated refresh block not found");
@@ -339,6 +347,13 @@ test("editor-poc.html reloads Forge buffer on parent-reload", () => {
assert.match(handler[1], /loadFile\(\);/);
});
test("editor-poc.html inserts assistant text on parent-insert (BUG-057)", () => {
assert.match(forgeSrc, /if \(e\.data\.type === 'parent-insert' && typeof e\.data\.text === 'string'\) \{/);
const handler = forgeSrc.match(/if \(e\.data\.type === 'parent-insert' && typeof e\.data\.text === 'string'\) \{([\s\S]*?)\n \}/);
assert.ok(handler, "parent-insert handler not found");
assert.match(handler[1], /insertAtCursor\(/);
});
test("style.css lets the inline editor fill the content area", () => {
assert.match(cssSrc, /\.editor-modal\.editor-inline-mode \{/);
assert.match(cssSrc, /\.editor-modal\.editor-inline-mode \{[\s\S]*?pointer-events: none;/);
@@ -365,6 +380,19 @@ test("style.css: #editor-body scrolls through the CodeMirror scroller only", ()
"global .cm-scroller min-height override reintroduces the double scrollbar");
});
// ── BUG-058: the line-number gutter follows the active theme ──
test("style.css themes the CodeMirror line-number gutter with CSS variables", () => {
const gutter = cssSrc.match(/\.cm-editor \.cm-gutters \{([^}]*)\}/);
assert.ok(gutter, "themed .cm-gutters rule not found");
assert.match(gutter[1], /background:\s*color-mix\([^;]*var\(--text-primary\)/,
"gutter background must derive from the theme, not CodeMirror's hardcoded #f5f5f5");
assert.match(gutter[1], /color:\s*var\(--text-secondary\)/);
assert.match(gutter[1], /border-right:\s*1px solid var\(--border\)/);
const active = cssSrc.match(/\.cm-editor \.cm-gutters \.cm-activeLineGutter \{([^}]*)\}/);
assert.ok(active, "themed .cm-activeLineGutter rule not found");
assert.match(active[1], /color-mix\([^;]*var\(--text-primary\)/);
});
// ── BUG-054: the shared save button must not stay stuck on the spinner ──
test("utils.js resetSaveButton restores the checkmark and re-enables the button", () => {
const fn = utilsSrc.match(/function resetSaveButton\(\) \{([\s\S]*?)\n\}/);
+8
View File
@@ -221,6 +221,14 @@ test("Forge autosave keeps the buffer dirty if edits arrive during the request",
assert.match(fn[1], /if \(val\(\) !== content\) \{ scheduleAutoSave\(\); return; \}/);
});
test("Forge leaves fullscreen before opening the shared assistant", () => {
const fn = FORGE_HTML.match(/function openAssistant\(\) \{([\s\S]*?)\n \}/);
assert.ok(fn, "openAssistant not found");
assert.match(fn[1], /document\.fullscreenElement/);
assert.match(fn[1], /document\.exitFullscreen\(\)/);
assert.match(fn[1], /postMessage\(\{ type: 'forge-open-ai' \}, '\*'\)/);
});
// ── Summary ────────────────────────────────────────────────────────────────
console.log(`\n${testCount} passed, ${failCount} failed`);
process.exit(failCount > 0 ? 1 : 0);
+95
View File
@@ -0,0 +1,95 @@
#!/usr/bin/env node
/**
* ObsiGate — Viewer PDF non-regression tests (BUG-060).
*
* Static checks on the source of the PDF inline viewer:
* - BUG-060 : the CSP set by BUG-034 puts `object-src 'none'`, which blocks
* `<embed>`/`<object>`. The PDF body was therefore never rendered (blank
* pages). The viewer must now use an `<iframe>` — permitted by
* `frame-src 'self'` since the PDF stream URL is same-origin.
*
* Usage: node tests/frontend/pdf-viewer.test.mjs
*/
import { strict as assert } from "node:assert";
import { readFileSync } from "node:fs";
import path from "node:path";
import { fileURLToPath } from "node:url";
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const ROOT = path.join(__dirname, "..", "..");
const viewer = readFileSync(path.join(ROOT, "frontend", "js", "viewer.js"), "utf8");
const main = readFileSync(path.join(ROOT, "backend", "main.py"), "utf8");
const css = readFileSync(path.join(ROOT, "frontend", "style.css"), "utf8");
function test(label, fn) {
try {
fn();
console.log(" \u2713 " + label);
} catch (err) {
console.error(" \u2717 " + label + "\n " + err.message);
process.exitCode = 1;
}
}
// ── viewer.js : le rendu PDF passe par une iframe ─────────────────────────
test("viewer.js — PDF branch renders the stream in an <iframe class=pdf-iframe>", () => {
const block = viewer.match(/if \(data\.is_pdf\) \{([\s\S]*?)\n \}/);
assert.ok(block, "PDF render block not found");
assert.match(
block[1],
/<iframe src="\$\{pdfUrl\}" class="pdf-iframe"/,
"PDF must use <iframe>, not <embed>/<object> (CSP object-src 'none' otherwise blocks it)",
);
assert.match(
block[1],
/\.pdf-iframe[\s\S]*?src="\$\{pdfUrl\}"/,
"iframe src must come from the /pdf/stream URL",
);
});
test("viewer.js — no <embed>/<object> left in the source", () => {
assert.doesNotMatch(viewer, /<embed\b/i, "<embed> is blocked by CSP object-src 'none'");
assert.doesNotMatch(viewer, /<object\b/i, "<object> is blocked by CSP object-src 'none'");
});
test("viewer.js — TOC still targets the pdf-iframe via contentWindow", () => {
assert.match(
viewer,
/document\.querySelector\('\.pdf-iframe'\)\.contentWindow\.location\.hash='page=\$\{item\.page\}'/,
"TOC links must keep navigating the iframe",
);
});
// ── backend : la CSP autorise le cadre same-origin ─────────────────────────
test("backend CSP — frame-src 'self' allows same-origin iframes", () => {
const csp = main.match(/frame-src ([^";]+);/);
assert.ok(csp, "CSP frame-src directive not found");
assert.ok(
csp[1].includes("'self'"),
`frame-src must allow 'self' (found: ${csp[1]}) — otherwise the PDF iframe is blocked`,
);
});
test("backend CSP — object-src 'none' stays in place (no defusing)", () => {
assert.match(
main,
/object-src 'none';/,
"object-src must stay locked to 'none': the fix is moving to <iframe>, not weakening CSP",
);
});
// ── style.css : l'iframe garde une hauteur utile ───────────────────────────
test("style.css — .pdf-iframe fills the viewer body", () => {
const rule = css.match(/\.pdf-iframe\s*\{([^}]*)\}/);
assert.ok(rule, ".pdf-iframe rule not found");
assert.match(rule[1], /flex:\s*1/, "iframe must stretch to fill the available height");
assert.match(rule[1], /min-height/, "iframe must keep its minimum height");
});
if (process.exitCode) {
console.error("\nPDF viewer tests FAILED");
} else {
console.log("\nAll PDF viewer tests passed.");
}
+27
View File
@@ -743,6 +743,33 @@ class TestSecretRedactor:
result = redact_file_content("hello world this is safe")
assert result == "hello world this is safe"
def test_git_sha_not_redacted(self):
"""BUG-035: a bare git commit SHA must not be mangled."""
from backend.secret_redactor import redact_file_content
sha = "a1b2c3d4e5f60718293a4b5c6d7e8f9012345678"
text = f"commit {sha}\nMerge: {sha}"
assert redact_file_content(text) == text
def test_sha256_checksum_not_redacted(self):
from backend.secret_redactor import redact_file_content
digest = "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
text = f"sha256:{digest} file.tar.gz"
assert redact_file_content(text) == text
def test_hex_secret_in_context_is_redacted(self):
from backend.secret_redactor import redact_file_content
secret = "0123456789abcdef0123456789abcdef01234567"
result = redact_file_content(f"api_key={secret}")
assert secret not in result
assert "MASQUÉ" in result
def test_ambiguous_bare_hex_left_intact(self):
"""A 40-char hex with no secret/hash keyword stays untouched."""
from backend.secret_redactor import redact_file_content
blob = "deadbeefdeadbeefdeadbeefdeadbeefdeadbeef"
text = f"value {blob} end"
assert redact_file_content(text) == text
# ═══════════════════════════════════════════════════════════════════
# Static / PWA caching (Cloudflare / mobile freshness)
+74
View File
@@ -44,6 +44,22 @@ class TestPasswordHashing:
result = hash_password("ab")
assert result is not None
def test_argon2_memory_recalibrated(self):
"""BUG-038: memory cost must stay at the OWASP 19 MiB recommendation."""
from backend.auth.password import (
ARGON2_MEMORY_COST_KIB,
ARGON2_PARALLELISM,
ARGON2_TIME_COST,
ph,
)
assert ARGON2_MEMORY_COST_KIB == 19456
assert ARGON2_TIME_COST == 2
assert ARGON2_PARALLELISM == 1
assert ph.memory_cost == 19456
assert ph.time_cost == 2
assert ph.parallelism == 1
# ═══════════════════════════════════════════════════════════════════
# JWT Handler
@@ -279,3 +295,61 @@ class TestMiddleware:
assert check_vault_access("Vault1", user) is True
assert check_vault_access("Vault3", user) is False
assert check_vault_access("Vault1", nobody) is False
# ═══════════════════════════════════════════════════════════════════
# Insecure (auth-disabled) deployment guard — BUG-037
# ═══════════════════════════════════════════════════════════════════
class TestInsecureAuthGuard:
def test_is_loopback_host(self):
from backend.auth.middleware import is_loopback_host
assert is_loopback_host(None) is True
assert is_loopback_host("127.0.0.1") is True
assert is_loopback_host("::1") is True
assert is_loopback_host("[::1]") is True
assert is_loopback_host("localhost") is True
assert is_loopback_host("0.0.0.0") is False
assert is_loopback_host("192.168.1.10") is False
def test_bind_host_from_argv(self):
from backend.auth.middleware import bind_host_from_argv
assert bind_host_from_argv(
["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8080"]
) == "0.0.0.0"
assert bind_host_from_argv(["uvicorn", "app", "--host=127.0.0.1"]) == "127.0.0.1"
assert bind_host_from_argv(["uvicorn", "app"]) is None
def test_guard_refuses_public_bind_without_optin(self, monkeypatch):
from backend import main
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
monkeypatch.delenv("OBSIGATE_ALLOW_INSECURE", raising=False)
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
with pytest.raises(RuntimeError):
main._guard_insecure_auth()
def test_guard_allows_loopback(self, monkeypatch):
from backend import main
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
monkeypatch.delenv("OBSIGATE_ALLOW_INSECURE", raising=False)
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "127.0.0.1"])
main._guard_insecure_auth() # must not raise
def test_guard_allows_explicit_optin(self, monkeypatch):
from backend import main
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "false")
monkeypatch.setenv("OBSIGATE_ALLOW_INSECURE", "true")
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
main._guard_insecure_auth() # must not raise
def test_guard_noop_when_auth_enabled(self, monkeypatch):
from backend import main
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
monkeypatch.setattr("sys.argv", ["uvicorn", "backend.main:app", "--host", "0.0.0.0"])
main._guard_insecure_auth() # must not raise
+36
View File
@@ -137,6 +137,42 @@ class TestLogin:
})
assert resp.status_code == 401
def test_locked_account_returns_401_not_429(self, auth_client, monkeypatch):
"""BUG-039: a locked account must be indistinguishable from an unknown one."""
import backend.auth.router as auth_router
monkeypatch.setattr(auth_router, "is_locked", lambda username: True)
resp = auth_client.post("/api/auth/login", json={
"username": "admin",
"password": "chab30",
})
assert resp.status_code == 401
assert "verrouill" not in resp.json()["detail"].lower()
def test_account_rate_limited_returns_401(self, auth_client, monkeypatch):
"""BUG-039: per-account throttling must not reveal the account exists."""
import backend.auth.router as auth_router
monkeypatch.setattr(auth_router, "is_account_rate_limited", lambda username: True)
resp = auth_client.post("/api/auth/login", json={
"username": "admin",
"password": "chab30",
})
assert resp.status_code == 401
def test_inactive_account_returns_401(self, auth_client, monkeypatch):
"""BUG-039: a disabled account answers like an unknown user."""
import backend.auth.router as auth_router
monkeypatch.setattr(auth_router, "get_user", lambda username: {
"username": username, "active": False, "password_hash": "x",
})
resp = auth_client.post("/api/auth/login", json={
"username": "admin",
"password": "chab30",
})
assert resp.status_code == 401
def test_login_remember_me(self, auth_client):
resp = auth_client.post("/api/auth/login", json={
"username": "admin",
+51
View File
@@ -73,6 +73,36 @@ def test_authenticate_websocket_invalid_token_returns_none(monkeypatch):
assert authenticate_websocket(_StubWebSocket(cookies={"access_token": "garbage"})) is None
def test_authenticate_websocket_query_token_rejected(monkeypatch):
"""BUG-036: the access token must never be accepted from the query string."""
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
from backend.auth.jwt_handler import create_access_token
token = create_access_token({
"username": "u", "role": "user", "vaults": ["*"], "display_name": "U",
})
ws = _StubWebSocket(query={"token": token})
assert authenticate_websocket(ws) is None
def test_authenticate_websocket_cookie_token_accepted(monkeypatch):
"""The HttpOnly access_token cookie remains the supported transport."""
monkeypatch.setenv("OBSIGATE_AUTH_ENABLED", "true")
import backend.auth.user_store as user_store
from backend.auth.jwt_handler import create_access_token
token = create_access_token({
"username": "u", "role": "user", "vaults": ["*"], "display_name": "U",
})
monkeypatch.setattr(user_store, "get_user", lambda username: {
"username": username, "role": "user", "vaults": ["*"],
"display_name": "U", "active": True,
})
user = authenticate_websocket(_StubWebSocket(cookies={"access_token": token}))
assert user is not None
assert user["username"] == "u"
# ---------------------------------------------------------------------------
# Manager unit tests (no WebSocket transport)
# ---------------------------------------------------------------------------
@@ -127,6 +157,27 @@ async def test_on_message_rejects_oversized_update(tmp_path: Path):
assert room.updates == []
@pytest.mark.asyncio
async def test_on_message_rejects_oversized_raw(tmp_path: Path):
"""BUG-036: oversized raw frames are dropped before parsing."""
target = tmp_path / "note.md"
target.write_text("x", encoding="utf-8")
manager = _make_manager(tmp_path)
room = CollabRoom(vault="V", path="note.md", file_path=target)
client = _FakeClient(conn_id=1)
import backend.collab as collab_mod
original = collab_mod.MAX_MESSAGE_CHARS
try:
collab_mod.MAX_MESSAGE_CHARS = 10
await manager._on_message(room, client, json.dumps({"type": "text", "text": "hello"}))
finally:
collab_mod.MAX_MESSAGE_CHARS = original
assert room.pending_text is None
class _FakeWebSocket:
def __init__(self):
self.sent: list[dict] = []
+230
View File
@@ -0,0 +1,230 @@
"""Unit tests for the connected sources (#92): Gitea & GitHub tools.
All HTTP calls are mocked (httpx.request monkeypatched) — the CI never talks
to a real Gitea/GitHub instance.
"""
import base64
from typing import Any
import pytest
import backend.tools.connected as connected
from backend.tools.context import ToolContext, ToolError, ToolMode
from backend.tools.registry import get_tool
class FakeResponse:
def __init__(self, json_data: Any = None, status_code: int = 200):
self._json = json_data
self.status_code = status_code
def json(self):
return self._json
def raise_for_status(self):
if self.status_code >= 400:
import httpx
raise httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
@pytest.fixture
def gitea_env(monkeypatch):
monkeypatch.setenv("OBSIGATE_GITEA_URL", "https://git.example.net")
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "tok-gitea")
@pytest.fixture
def github_env(monkeypatch):
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "tok-gh")
def _patch_request(monkeypatch, handler):
def fake_request(method, url, headers=None, timeout=None, follow_redirects=False, **kw):
captured = {"method": method, "url": str(url), "headers": headers or {},
"params": kw.get("params")}
return handler(captured)
monkeypatch.setattr(connected.httpx, "request", fake_request)
class TestRegistration:
@pytest.mark.parametrize("name", ["git_list_repos", "git_search_issues", "git_get_file"])
def test_tools_registered_read(self, name):
from backend.tools.context import ToolRisk
spec = get_tool(name)
assert spec is not None
assert spec.risk == ToolRisk.READ
assert "gitea" in spec.description or "github" in spec.description.lower()
class TestConfiguration:
def test_gitea_requires_base_url(self, monkeypatch):
monkeypatch.delenv("OBSIGATE_GITEA_URL", raising=False)
with pytest.raises(ToolError) as ei:
connected._provider_base("gitea")
assert ei.value.code == "provider_not_configured"
def test_unknown_provider_rejected(self):
with pytest.raises(ToolError) as ei:
connected._provider_base("gitlab")
assert ei.value.code == "invalid_arguments"
def test_gitea_token_sent_as_header(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"data": []})
_patch_request(monkeypatch, handler)
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert captured["headers"]["Authorization"] == "token tok-gitea"
class TestListRepos:
def test_gitea_search_endpoint(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"data": [
{"name": "ObsiGate", "full_name": "bruno/ObsiGate",
"html_url": "https://git.example.net/bruno/ObsiGate",
"description": "vault gateway", "updated_at": "2026-09-01",
"private": False},
]})
_patch_request(monkeypatch, handler)
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert "/api/v1/repos/search" in captured["url"]
assert out["repos"][0]["name"] == "ObsiGate"
assert out["count"] == 1
def test_github_single_repo(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={
"name": "ObsiGate", "full_name": "bruno/ObsiGate",
"html_url": "https://github.com/bruno/ObsiGate",
"description": "", "updated_at": "2026-09-02", "private": True,
})
_patch_request(monkeypatch, handler)
out = connected.git_list_repos(_ctx(), connected.GitProviderInput(
provider="github", repo="bruno/ObsiGate"))
assert captured["url"].endswith("/repos/bruno/ObsiGate")
assert out["repos"][0]["full_name"] == "bruno/ObsiGate"
assert out["repos"][0]["private"] is True
class TestSearchIssues:
def test_gitea_scoped_to_repo(self, gitea_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data=[
{"number": 12, "title": "Bug affichage", "html_url": "https://x/12",
"state": "open"},
])
_patch_request(monkeypatch, handler)
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
provider="gitea", query="affichage", repo="bruno/ObsiGate"))
assert "/repos/bruno/ObsiGate/issues" in captured["url"]
assert out["issues"][0]["id"] == 12
assert out["issues"][0]["pull_request"] is False
def test_github_search_syntax(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={"items": [
{"number": 5, "title": "Crash on save", "html_url": "https://gh/5",
"state": "open", "pull_request": {"url": "x"}},
]})
_patch_request(monkeypatch, handler)
out = connected.git_search_issues(_ctx(), connected.GitSearchIssuesInput(
provider="github", query="crash", repo="bruno/ObsiGate", state="open"))
assert "/search/issues" in captured["url"]
assert "repo:bruno/ObsiGate" in captured["params"]["q"]
assert out["issues"][0]["pull_request"] is True
class TestGetFile:
def test_gitea_base64_content_decoded(self, gitea_env, monkeypatch):
payload = base64.b64encode("# Readme\n\nBonjour".encode()).decode()
def handler(_captured):
return FakeResponse(json_data={
"path": "README.md", "size": 15, "encoding": "base64", "content": payload,
})
_patch_request(monkeypatch, handler)
out = connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="gitea", repo="bruno/ObsiGate", path="README.md"))
assert "Bonjour" in out["content"]
assert out["truncated"] is False
def test_github_ref_parameter(self, github_env, monkeypatch):
captured = {}
def handler(captured_req):
captured.update(captured_req)
return FakeResponse(json_data={
"path": "a.md", "size": 1, "encoding": "base64",
"content": base64.b64encode(b"x").decode(),
})
_patch_request(monkeypatch, handler)
connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="github", repo="o/r", path="a.md", ref="v2.9.0"))
assert captured["url"].endswith("?ref=v2.9.0")
def test_missing_repo_or_path_rejected(self, gitea_env):
with pytest.raises(ToolError) as ei:
connected.git_get_file(_ctx(), connected.GitGetFileInput(
provider="gitea", repo="", path="a.md"))
assert ei.value.code == "invalid_arguments"
class TestErrors:
def test_404_maps_to_not_found(self, gitea_env, monkeypatch):
def handler(_captured):
return FakeResponse(json_data={}, status_code=404)
_patch_request(monkeypatch, handler)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "not_found"
def test_401_maps_to_permission_denied(self, gitea_env, monkeypatch):
def handler(_captured):
return FakeResponse(json_data={}, status_code=401)
_patch_request(monkeypatch, handler)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "permission_denied"
def test_network_error_maps_to_tool_error(self, gitea_env, monkeypatch):
import httpx
def fake_request(*a, **kw):
raise httpx.ConnectError("down")
monkeypatch.setattr(connected.httpx, "request", fake_request)
with pytest.raises(ToolError) as ei:
connected.git_list_repos(_ctx(), connected.GitProviderInput(provider="gitea"))
assert ei.value.code == "connected_source_unavailable"
+135
View File
@@ -0,0 +1,135 @@
"""Unit tests for the bounded site crawler (#92): crawl_site (WRITE + confirmation)."""
from typing import Any
import pytest
import backend.tools.crawler as crawler
from backend.tools.api import ToolConfirmationRequired, ToolContext, ToolError, call_tool
class FakeResponse:
def __init__(self, content: bytes = b"", status_code: int = 200,
headers: dict | None = None):
self.content = content
self.status_code = status_code
self.headers = headers or {"content-type": "text/html; charset=utf-8"}
self.encoding = "utf-8"
def raise_for_status(self):
pass
def _ctx() -> ToolContext:
return ToolContext(
user={"username": "tester", "role": "admin", "vaults": ["*"]},
audit_enabled=False,
)
@pytest.fixture
def vault(tmp_path, monkeypatch):
vault_dir = tmp_path / "Vault"
vault_dir.mkdir()
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
return vault_dir
PAGE_A = (
b"<html><head><title>Docs</title></head><body>"
b"<p>Bienvenue sur la documentation.</p>"
b'<a href="/page-b">Suite</a><a href="https://other.dev/x">ext</a>'
b"</body></html>"
)
PAGE_B = (
b"<html><head><title>Page B</title></head><body><p>Details ici.</p></body></html>"
)
@pytest.fixture
def local_urls(monkeypatch):
"""Skip the DNS-based SSRF guard: test hosts are fake, HTTP is mocked."""
monkeypatch.setattr(crawler, "_assert_public_http_url", lambda url: url)
@pytest.fixture
def two_pages(monkeypatch, local_urls):
def fake_get(url, **kw):
url = str(url)
if url.endswith("/page-b"):
return FakeResponse(content=PAGE_B)
return FakeResponse(content=PAGE_A)
monkeypatch.setattr(crawler.httpx, "get", fake_get)
class TestConfirmation:
def test_requires_confirmation(self, vault, two_pages):
with pytest.raises(ToolConfirmationRequired):
call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "Crawls/docs.md",
})
class TestCrawl:
def test_saves_same_host_pages(self, vault, two_pages):
out = call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "Crawls/docs.md",
}, confirm=True)
assert out.ok and out.data["pages"] == 2
digest = (vault / "Crawls" / "docs.md").read_text(encoding="utf-8")
assert "# Crawl de docs.example.dev" in digest
assert "Bienvenue sur la documentation." in digest
assert "Details ici." in digest
assert "other.dev" not in digest
def test_max_pages_bound(self, vault, monkeypatch, local_urls):
# A link farm: every page links to a new page — cap at max_pages.
def fake_get(url, **kw):
url = str(url)
n = int(url.rsplit("/", 1)[-1] or 0)
return FakeResponse(
content=f"<html><head><title>P{n}</title></head><body>"
f"<p>page {n}</p><a href=\"/{n + 1}\">next</a></body></html>".encode())
monkeypatch.setattr(crawler.httpx, "get", fake_get)
out = call_tool("crawl_site", _ctx(), {
"url": "https://farm.example.dev/0",
"vault": "Vault", "path": "farm.md", "max_pages": 3,
}, confirm=True)
assert out.ok and out.data["pages"] == 3
def test_no_pages_recovered(self, vault, monkeypatch, local_urls):
def dead_get(*a, **kw):
raise crawler.httpx.ConnectError("down")
monkeypatch.setattr(crawler.httpx, "get", dead_get)
with pytest.raises(ToolError) as ei:
call_tool("crawl_site", _ctx(), {
"url": "https://dead.example.dev/", "vault": "Vault", "path": "x.md",
}, confirm=True)
assert ei.value.code == "crawl_failed"
def test_internal_url_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("crawl_site", _ctx(), {
"url": "http://127.0.0.1:8080/", "vault": "Vault", "path": "x.md",
}, confirm=True)
assert ei.value.code in ("ssrf_blocked", "dns_error")
def test_binary_content_skipped(self, vault, monkeypatch, local_urls):
def fake_get(url, **kw):
url = str(url)
if url.endswith("/x.pdf"):
return FakeResponse(content=b"%PDF-1.4", headers={"content-type": "application/pdf"})
return FakeResponse(content=PAGE_A)
monkeypatch.setattr(crawler.httpx, "get", fake_get)
out = call_tool("crawl_site", _ctx(), {
"url": "https://docs.example.dev/start",
"vault": "Vault", "path": "docs.md", "max_pages": 5,
}, confirm=True)
assert out.ok and out.data["pages"] >= 1
+197
View File
@@ -0,0 +1,197 @@
"""Unit tests for the document-production tools (#92): create_xlsx, create_docx,
create_csv, create_pdf — WRITE risk, confirmation gating, vault persistence."""
import csv
import io
from pathlib import Path
import pytest
from backend.tools.api import (
ToolConfirmationRequired,
ToolContext,
ToolError,
call_tool,
get_tool,
)
from backend.tools.context import ToolRisk
@pytest.fixture
def vault(tmp_path, monkeypatch):
"""A minimal configured vault (index entry patched, no full build)."""
vault_dir = tmp_path / "Vault"
vault_dir.mkdir()
monkeypatch.setitem(__import__("backend.indexer", fromlist=["index"]).index,
"Vault", {"name": "Vault", "path": str(vault_dir), "config": {}})
return vault_dir
def _ctx() -> ToolContext:
return ToolContext(
user={"username": "tester", "role": "admin", "vaults": ["*"]},
audit_enabled=False,
)
class TestRegistry:
@pytest.mark.parametrize("name", ["create_xlsx", "create_docx", "create_csv", "create_pdf"])
def test_write_risk_and_confirmation(self, name):
spec = get_tool(name)
assert spec is not None
assert spec.risk == ToolRisk.WRITE
assert spec.requires_confirmation is True
class TestConfirmationGating:
def test_csv_requires_confirmation(self, vault):
with pytest.raises(ToolConfirmationRequired):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "data.csv",
"rows": [["a", "b"], [1, 2]],
})
def test_pdf_requires_confirmation(self, vault):
with pytest.raises(ToolConfirmationRequired):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "doc.pdf", "title": "T", "content": "# H\npara",
})
class TestCreateCsv:
def test_creates_file_in_vault(self, vault):
out = call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "Exports/data.csv",
"rows": [["nom", "score"], ["alice", 12], ["bob", 9.5]],
}, confirm=True)
assert out.ok and out.data["success"] is True
path = vault / "Exports" / "data.csv"
assert path.exists()
rows = list(csv.reader(io.StringIO(path.read_text(encoding="utf-8"))))
assert rows[0] == ["nom", "score"]
assert rows[2] == ["bob", "9.5"]
def test_semicolon_delimiter(self, vault):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.csv", "delimiter": ";",
"rows": [["a", "b"], [1, 2]],
}, confirm=True)
content = (vault / "d.csv").read_text(encoding="utf-8")
assert "a;b" in content
def test_wrong_extension_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.txt", "rows": [["a"], [1]],
}, confirm=True)
assert ei.value.code == "invalid_arguments"
def test_empty_rows_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "d.csv", "rows": [],
}, confirm=True)
class TestCreateXlsx:
def test_creates_readable_workbook(self, vault):
call_tool("create_xlsx", _ctx(), {
"vault": "Vault", "path": "Rapports/budget.xlsx",
"rows": [["item", "cout"], ["serveur", 1200], ["licence", 300]],
"sheet_name": "Budget",
}, confirm=True)
from openpyxl import load_workbook
wb = load_workbook(vault / "Rapports" / "budget.xlsx")
ws = wb.active
assert ws.title == "Budget"
assert ws.cell(row=1, column=1).value == "item"
assert ws.cell(row=2, column=2).value == 1200
def test_wrong_extension_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_xlsx", _ctx(), {
"vault": "Vault", "path": "b.docx", "rows": [["a"], [1]],
}, confirm=True)
class TestCreateDocx:
def test_creates_readable_document(self, vault):
call_tool("create_docx", _ctx(), {
"vault": "Vault", "path": "rapport.docx",
"title": "Rapport hebdo", "paragraphs": ["Premier point.", "Second point."],
}, confirm=True)
import docx as docx_lib
doc = docx_lib.Document(str(vault / "rapport.docx"))
texts = [p.text for p in doc.paragraphs]
assert "Rapport hebdo" in texts
assert "Second point." in texts
def test_no_paragraphs_rejected(self, vault):
with pytest.raises(ToolError):
call_tool("create_docx", _ctx(), {
"vault": "Vault", "path": "r.docx", "paragraphs": [],
}, confirm=True)
class TestCreatePdf:
def test_creates_valid_pdf(self, vault):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "docs/archi.pdf",
"title": "Architecture", "content": "# Titre\n\nUn paragraphe.\n## Sous-titre\nAutre texte.",
}, confirm=True)
raw = (vault / "docs" / "archi.pdf").read_bytes()
assert raw.startswith(b"%PDF")
def test_long_content_truncated(self, vault):
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "big.pdf", "title": "T", "content": "x" * 500_000,
}, confirm=True)
assert (vault / "big.pdf").exists()
def test_markdown_tables_go_through_the_export_pipeline(self, vault, monkeypatch):
# The document-page pipeline (mistune tables + WeasyPrint) must be
# used when available: capture the HTML handed to the PDF generator.
import sys
import types
captured = {}
fake = types.ModuleType("backend.pdf_export")
def fake_build(html, title, **kw):
captured["html"] = html
captured["title"] = title
return "<html>" + html + "</html>"
fake.build_pdf_html = fake_build
fake.generate_pdf = lambda html, title=None, **kw: b"%PDF-fake"
monkeypatch.setitem(sys.modules, "backend.pdf_export", fake)
out = call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "table.pdf", "title": "Rapport",
"content": "# T\n\n| a | b |\n|---|---|\n| 1 | 2 |",
}, confirm=True)
assert out.ok
assert "<table>" in captured["html"]
assert captured["title"] == "Rapport"
assert (vault / "table.pdf").read_bytes() == b"%PDF-fake"
def test_reportlab_fallback_when_weasyprint_missing(self, vault, monkeypatch):
# sys.modules[name] = None makes `from backend.pdf_export import …`
# raise ImportError → the simplified renderer must take over.
import sys
monkeypatch.setitem(sys.modules, "backend.pdf_export", None)
call_tool("create_pdf", _ctx(), {
"vault": "Vault", "path": "fb.pdf", "title": "T", "content": "# H\ntexte",
}, confirm=True)
assert (vault / "fb.pdf").read_bytes().startswith(b"%PDF")
class TestVaultSafety:
def test_path_outside_vault_rejected(self, vault):
with pytest.raises(ToolError) as ei:
call_tool("create_csv", _ctx(), {
"vault": "Vault", "path": "../outside.csv", "rows": [["a"], [1]],
}, confirm=True)
assert ei.value.code in ("path_outside_vault", "invalid_arguments", "tool_execution_error")
+90 -3
View File
@@ -275,6 +275,7 @@ class TestPdfIndexing:
"""When a vault directory is scanned with .pdf files, they appear in files list.
Uses the public _scan_vault() helper directly — no global state needed.
BUG-040: text extraction is deferred, so the scan only carries metadata.
"""
from backend.indexer import _scan_vault
@@ -286,10 +287,96 @@ class TestPdfIndexing:
names = {f["path"] for f in result["files"]}
assert "a.pdf" in names
assert "b.pdf" in names
# The PDF content should have been extracted.
# The scan defers the expensive text extraction.
a_file = next(f for f in result["files"] if f["path"] == "a.pdf")
assert "ObsiGate test PDF" in (a_file.get("content") or "")
assert "uniqueword0" in (a_file.get("content") or "")
assert a_file["content"] == ""
assert a_file["pdf_text_pending"] is True
class TestPdfLazyEnrichment:
"""BUG-040: PDF text is extracted in a deferred background pass."""
def test_scan_defers_pdf_text_extraction(self, pdf_dir: Path, tmp_path: Path):
from backend.indexer import _scan_vault
vault_root = tmp_path / "vault"
vault_root.mkdir()
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
result = _scan_vault("v", str(vault_root), {})
a_file = next(f for f in result["files"] if f["path"] == "a.pdf")
assert a_file["content"] == ""
assert a_file["content_preview"] == ""
assert a_file["pdf_text_pending"] is True
@pytest.mark.asyncio
async def test_enrich_pdf_texts_fills_content_and_clears_flag(
self, pdf_dir: Path, tmp_path: Path
):
import backend.indexer as idx
vault_root = tmp_path / "vault"
vault_root.mkdir()
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
file_info = {
"path": "a.pdf",
"title": "a",
"tags": [],
"content": "",
"content_preview": "",
"size": 0,
"modified": "",
"extension": ".pdf",
"pdf_text_pending": True,
}
with idx._index_lock:
idx.index["LazyV"] = {
"files": [file_info],
"tags": {},
"path": str(vault_root),
"paths": [],
}
try:
count = await idx.enrich_pdf_texts("LazyV")
assert count == 1
assert "ObsiGate test PDF" in file_info["content"]
assert "uniqueword0" in file_info["content"]
assert file_info["content_preview"]
assert "pdf_text_pending" not in file_info
finally:
with idx._index_lock:
idx.index.pop("LazyV", None)
@pytest.mark.asyncio
async def test_enrich_pdf_texts_skips_other_vaults(self, pdf_dir: Path, tmp_path: Path):
import backend.indexer as idx
vault_root = tmp_path / "vault"
vault_root.mkdir()
shutil.copy2(pdf_dir / "simple.pdf", vault_root / "a.pdf")
file_info = {
"path": "a.pdf",
"title": "a",
"tags": [],
"content": "",
"content_preview": "",
"size": 0,
"modified": "",
"extension": ".pdf",
"pdf_text_pending": True,
}
with idx._index_lock:
idx.index["OtherV"] = {
"files": [file_info],
"tags": {},
"path": str(vault_root),
"paths": [],
}
try:
assert await idx.enrich_pdf_texts("LazyV") == 0
assert file_info["content"] == ""
finally:
with idx._index_lock:
idx.index.pop("OtherV", None)
# ── Search filter `ext:` ───────────────────────────────────────────────────
+163
View File
@@ -0,0 +1,163 @@
"""Unit tests for the tool/connected-source key store (#103).
Covers ``backend.tools.secrets`` (precedence, masking, whitelist) and the
``/api/config/tool-keys`` endpoints (masked GET, POST, DELETE, admin-only).
"""
import json
from pathlib import Path
import pytest
from backend.tools.secrets import (
TOOL_KEY_NAMES,
delete_tool_key,
get_tool_key,
is_secret_name,
mask_value,
set_tool_key,
)
@pytest.fixture
def key_store(tmp_path, monkeypatch):
"""Isolated key store directory."""
monkeypatch.setenv("OBSIGATE_DATA_DIR", str(tmp_path))
yield tmp_path
def _write_store(tmp_path, data):
(tmp_path / "api_keys.json").write_text(json.dumps(data), encoding="utf-8")
class TestGetToolKey:
def test_stored_value_takes_precedence_over_env(self, key_store, monkeypatch):
_write_store(key_store, {"OBSIGATE_GITHUB_TOKEN": "stored-token"})
monkeypatch.setenv("OBSIGATE_GITHUB_TOKEN", "env-token")
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == "stored-token"
def test_env_fallback_when_not_stored(self, key_store, monkeypatch):
monkeypatch.setenv("OBSIGATE_GITEA_TOKEN", "env-token")
assert get_tool_key("OBSIGATE_GITEA_TOKEN") == "env-token"
def test_missing_everywhere_returns_empty(self, key_store, monkeypatch):
monkeypatch.delenv("OBSIGATE_GITHUB_TOKEN", raising=False)
assert get_tool_key("OBSIGATE_GITHUB_TOKEN") == ""
def test_non_whitelisted_name_uses_env_only(self, key_store, monkeypatch):
_write_store(key_store, {"OTHER_KEY": "stored"})
monkeypatch.setenv("OTHER_KEY", "env")
assert get_tool_key("OTHER_KEY") == "env"
def test_whitelist_covers_expected_names(self):
assert set(TOOL_KEY_NAMES) == {
"OBSIGATE_TAVILY_API_KEY",
"OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY",
"OBSIGATE_EXA_API_KEY",
"OBSIGATE_GITEA_URL",
"OBSIGATE_GITEA_TOKEN",
"OBSIGATE_GITHUB_TOKEN",
}
class TestSetDelete:
def test_set_then_get_roundtrip(self, key_store):
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == "tvly-1234"
def test_set_empty_value_deletes_entry(self, key_store):
set_tool_key("OBSIGATE_TAVILY_API_KEY", "tvly-1234")
set_tool_key("OBSIGATE_TAVILY_API_KEY", "")
assert get_tool_key("OBSIGATE_TAVILY_API_KEY") == ""
assert "OBSIGATE_TAVILY_API_KEY" not in json.loads(
(key_store / "api_keys.json").read_text(encoding="utf-8")
)
def test_delete_removes_and_reports(self, key_store):
set_tool_key("OBSIGATE_GITEA_URL", "https://git.example.net")
assert delete_tool_key("OBSIGATE_GITEA_URL") is True
assert delete_tool_key("OBSIGATE_GITEA_URL") is False
def test_unknown_name_rejected(self, key_store):
with pytest.raises(ValueError):
set_tool_key("NOT_WHITELISTED", "x")
with pytest.raises(ValueError):
delete_tool_key("NOT_WHITELISTED")
def test_store_keeps_other_entries(self, key_store):
_write_store(key_store, {"DEEPSEEK_API_KEY": "sk-existing"})
set_tool_key("OBSIGATE_EXA_API_KEY", "exa-key")
data = json.loads((key_store / "api_keys.json").read_text(encoding="utf-8"))
assert data["DEEPSEEK_API_KEY"] == "sk-existing"
assert data["OBSIGATE_EXA_API_KEY"] == "exa-key"
class TestMasking:
def test_urls_returned_clear(self):
assert mask_value("OBSIGATE_GITEA_URL", "https://git.example.net") == \
"https://git.example.net"
def test_tokens_masked(self):
masked = mask_value("OBSIGATE_GITHUB_TOKEN", "ghp_abcdefgh1234")
assert masked.startswith("ghp_")
assert "abcdefgh1234" not in masked
assert "..." in masked
def test_short_secret_fully_masked(self):
assert mask_value("OBSIGATE_EXA_API_KEY", "abc") == "***"
def test_secret_detection(self):
assert is_secret_name("OBSIGATE_GITEA_TOKEN")
assert is_secret_name("OBSIGATE_TAVILY_API_KEY")
assert not is_secret_name("OBSIGATE_GITEA_URL")
class TestToolKeysAPI:
def _login(self, admin_client):
resp = admin_client.post(
"/api/auth/login", json={"username": "admin", "password": "chab30"}
)
assert resp.status_code == 200, resp.text
token = resp.json()["access_token"]
return {"Authorization": f"Bearer {token}"}
def test_get_masks_tokens_shows_urls(self, admin_client, key_store):
headers = self._login(admin_client)
_write_store(key_store, {
"OBSIGATE_GITEA_URL": "https://git.example.net",
"OBSIGATE_GITHUB_TOKEN": "ghp_abcdefgh1234",
})
resp = admin_client.get("/api/config/tool-keys", headers=headers)
assert resp.status_code == 200
data = resp.json()
assert data["OBSIGATE_GITEA_URL"] == "https://git.example.net"
assert "abcdefgh1234" not in data["OBSIGATE_GITHUB_TOKEN"]
assert data["OBSIGATE_GITHUB_TOKEN"].startswith("ghp_")
def test_post_roundtrip_then_delete(self, admin_client, key_store):
headers = self._login(admin_client)
resp = admin_client.post(
"/api/config/tool-keys",
headers=headers,
json={"OBSIGATE_EXA_API_KEY": "exa-key-1234"},
)
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
assert get_tool_key("OBSIGATE_EXA_API_KEY") == "exa-key-1234"
resp = admin_client.delete("/api/config/tool-keys/OBSIGATE_EXA_API_KEY", headers=headers)
assert resp.status_code == 200
assert resp.json()["status"] == "deleted"
assert get_tool_key("OBSIGATE_EXA_API_KEY") == ""
def test_post_unknown_name_rejected(self, admin_client, key_store):
headers = self._login(admin_client)
resp = admin_client.post(
"/api/config/tool-keys", headers=headers, json={"MY_SECRET": "x"}
)
assert resp.status_code == 400
def test_requires_admin(self, admin_client, key_store):
resp = admin_client.get("/api/config/tool-keys", headers={})
assert resp.status_code in (401, 403)
+82
View File
@@ -0,0 +1,82 @@
"""Unit tests for the SQLite web cache (backend.tools.webcache, #92)."""
import pytest
from backend.tools import webcache
@pytest.fixture
def cache_enabled(tmp_path, monkeypatch):
"""Enable the cache against an isolated file with a short TTL."""
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
webcache._schema_ready = False
yield
webcache._schema_ready = False
class TestCacheKey:
def test_deterministic_and_payload_sensitive(self):
a = webcache.cache_key("search", {"q": "pizza", "page": 1})
b = webcache.cache_key("search", {"page": 1, "q": "pizza"})
c = webcache.cache_key("search", {"q": "pasta", "page": 1})
d = webcache.cache_key("fetch", {"q": "pizza", "page": 1})
assert a == b
assert a != c
assert a != d
class TestCacheRoundTrip:
def test_set_get_roundtrip(self, cache_enabled):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"results": [1, 2, 3], "provider": "tavily"})
assert webcache.cache_get(key) == {"results": [1, 2, 3], "provider": "tavily"}
def test_miss_returns_none(self, cache_enabled):
assert webcache.cache_get("search:unknown") is None
def test_overwrite_updates_value(self, cache_enabled):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
webcache.cache_set(key, {"v": 2})
assert webcache.cache_get(key) == {"v": 2}
def test_ttl_expiry(self, tmp_path, monkeypatch):
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path / "cache.sqlite3"))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
webcache._schema_ready = False
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
assert webcache.cache_get(key) == {"v": 1}
import time
time.sleep(0.15)
assert webcache.cache_get(key) is None
def test_disabled_when_ttl_zero(self, cache_enabled, monkeypatch):
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0")
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
assert webcache.cache_get(key) is None
def test_purge_expired(self, cache_enabled, monkeypatch):
key = webcache.cache_key("search", {"q": "x"})
webcache.cache_set(key, {"v": 1})
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "0.05")
import time
time.sleep(0.15)
assert webcache.purge_expired() >= 1
assert webcache.cache_get(key) is None
def test_corrupt_db_degrades_silently(self, tmp_path, monkeypatch):
# A directory as cache file breaks sqlite3.connect: the cache must
# disable itself instead of breaking the tools.
monkeypatch.setenv("OBSIGATE_WEB_CACHE_PATH", str(tmp_path))
monkeypatch.setenv("OBSIGATE_WEB_CACHE_TTL", "60")
webcache._schema_ready = False
try:
assert webcache.cache_get("search:x") is None
webcache.cache_set("search:x", {"v": 1})
finally:
webcache._schema_ready = False
+155
View File
@@ -0,0 +1,155 @@
"""Unit tests for the keyed web-search providers (#92): Tavily, Brave,
SerpAPI, Exa — plus provider ordering and the transient retry."""
import httpx
import pytest
import backend.tools.web as web
from backend.tools.context import ToolContext, ToolError, ToolMode
class FakeResponse:
def __init__(self, json_data=None, status_code=200, content=b""):
self._json = json_data
self.status_code = status_code
self.content = content
self.encoding = "utf-8"
self.headers = {}
def json(self):
return self._json
def raise_for_status(self):
if self.status_code >= 400:
raise web.httpx.HTTPStatusError("boom", request=None, response=self) # type: ignore[arg-type]
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
def _no_fallback(monkeypatch):
"""Limit the chain to the provider under test (no searxng/ddg/bing noise)."""
monkeypatch.setattr(web, "WEB_FALLBACK_ENABLED", False)
monkeypatch.setattr(web, "SEARXNG_URL", "http://searxng.invalid")
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: (_ for _ in ()).throw(
httpx.ConnectError("offline")))
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: (_ for _ in ()).throw(
httpx.ConnectError("offline")))
class TestKeyedProviderParsers:
def test_tavily_maps_results(self, monkeypatch):
captured = {}
def fake_post(url, json=None, **kw):
captured["url"] = url
captured["payload"] = json
return FakeResponse(json_data={"results": [
{"title": "T", "url": "https://a.dev", "content": "c" * 800},
]})
monkeypatch.setattr(web.httpx, "post", fake_post)
results, engines = web._search_tavily("q", web.WebSearchInput(query="q", max_results=3))
assert results[0]["title"] == "T"
assert len(results[0]["snippet"]) <= 600
assert engines == []
assert captured["payload"]["api_key"] == ""
assert captured["payload"]["max_results"] == 3
def test_brave_maps_results(self, monkeypatch):
captured = {}
def fake_get(url, params=None, headers=None, **kw):
captured["url"] = url
captured["headers"] = headers
return FakeResponse(json_data={"web": {"results": [
{"title": "B", "url": "https://b.dev", "description": "d"},
]}})
monkeypatch.setattr(web.httpx, "get", fake_get)
results, _ = web._search_brave("q", web.WebSearchInput(query="q"))
assert results[0]["title"] == "B"
assert "api.search.brave.com" in str(captured["url"])
def test_serpapi_maps_results(self, monkeypatch):
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
"organic_results": [{"title": "S", "link": "https://s.dev", "snippet": "sn"}],
}))
results, _ = web._search_serpapi("q", web.WebSearchInput(query="q"))
assert results[0]["url"] == "https://s.dev"
def test_exa_maps_results(self, monkeypatch):
monkeypatch.setattr(web.httpx, "post", lambda *a, **kw: FakeResponse(json_data={
"results": [{"title": "E", "url": "https://e.dev", "text": "t" * 900}],
}))
results, _ = web._search_exa("q", web.WebSearchInput(query="q"))
assert results[0]["title"] == "E"
assert len(results[0]["snippet"]) <= 600
class TestProviderChain:
def test_no_key_falls_back_to_searxng(self, monkeypatch):
for var in ("OBSIGATE_TAVILY_API_KEY", "OBSIGATE_BRAVE_API_KEY",
"OBSIGATE_SERPAPI_API_KEY", "OBSIGATE_EXA_API_KEY"):
monkeypatch.delenv(var, raising=False)
chain = [name for name, _ in web._provider_chain()]
assert chain[0] == "searxng"
assert "tavily" not in chain
def test_keyed_provider_used_first_when_key_set(self, monkeypatch):
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
chain = [name for name, _ in web._provider_chain()]
assert chain[0] == "tavily"
assert "brave" not in chain # no key → skipped
def test_explicit_order_env(self, monkeypatch):
monkeypatch.setenv("OBSIGATE_TAVILY_API_KEY", "k")
monkeypatch.setenv("OBSIGATE_EXA_API_KEY", "k")
monkeypatch.setenv("OBSIGATE_WEB_PROVIDERS", "exa,unknown,tavily")
chain = [name for name, _ in web._provider_chain()]
assert chain[:2] == ["exa", "tavily"]
def test_search_uses_keyed_provider_first(self, monkeypatch):
_no_fallback(monkeypatch)
monkeypatch.setenv("OBSIGATE_BRAVE_API_KEY", "k")
monkeypatch.setattr(web.httpx, "get", lambda *a, **kw: FakeResponse(json_data={
"web": {"results": [{"title": "B", "url": "https://b.dev", "description": "d"}]},
}))
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert out["provider"] == "brave"
assert out["count"] == 1
class TestRetry:
def test_transient_error_retried_then_succeeds(self, monkeypatch):
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
calls = {"n": 0}
def flaky_get(*a, **kw):
calls["n"] += 1
if calls["n"] == 1:
raise httpx.ConnectError("blip")
return FakeResponse(json_data={"results": [
{"title": "A", "url": "https://a.dev", "content": "x"}]})
monkeypatch.setattr(web.httpx, "get", flaky_get)
out = web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert out["provider"] == "searxng"
assert calls["n"] == 2
def test_persistent_error_not_retried_forever(self, monkeypatch):
monkeypatch.setattr(web, "WEB_RETRY_ATTEMPTS", 1)
monkeypatch.setattr(web, "_provider_chain", lambda: [("searxng", web._search_searxng)])
calls = {"n": 0}
def dead_get(*a, **kw):
calls["n"] += 1
raise httpx.ConnectError("down")
monkeypatch.setattr(web.httpx, "get", dead_get)
with pytest.raises(ToolError) as ei:
web.web_search(_ctx(), web.WebSearchInput(query="q"))
assert ei.value.code == "web_search_unavailable"
assert calls["n"] == 2 # initial + 1 retry, per provider
+73
View File
@@ -0,0 +1,73 @@
"""Unit tests for the dynamic rendering path (#92): fetch_url(render=True)."""
import pytest
import backend.tools.web as web
from backend.tools import webrender
from backend.tools.context import ToolContext, ToolError, ToolMode
from backend.tools.registry import get_tool
def _ctx() -> ToolContext:
return ToolContext(user={"username": "tester", "vaults": []}, mode=ToolMode.IN_APP)
class TestRegistration:
def test_render_param_exposed_in_schema(self):
spec = get_tool("fetch_url")
assert spec is not None
assert "render" in spec.input_model.model_fields
class TestRenderUnavailable:
def test_missing_playwright_clear_error(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: False)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert ei.value.code == "playwright_unavailable"
def test_ssrf_guard_applied_before_render(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="http://127.0.0.1:9222/devtools", render=True))
assert ei.value.code in ("ssrf_blocked", "dns_error")
class TestRenderSuccess:
def test_fetch_url_delegates_to_worker(self, monkeypatch):
captured = {}
def fake_render(url):
captured["url"] = url
return {"url": url, "status": 200, "title": "SPA",
"text": "dynamic content", "rendered": True, "truncated": False}
monkeypatch.setattr(webrender, "render_page", fake_render)
out = web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert captured["url"] == "https://example.com/spa"
assert out["rendered"] is True
assert "dynamic content" in out["text"]
def test_worker_failure_maps_to_tool_error(self, monkeypatch):
monkeypatch.setattr(webrender, "_playwright_available", lambda: True)
def boom(url):
raise RuntimeError("chromium crashed")
# The executor re-raises the worker exception on .result(); render_page
# must wrap it into a ToolError instead of leaking a bare exception.
monkeypatch.setattr(webrender, "_render_in_worker", boom)
with pytest.raises(ToolError) as ei:
web.fetch_url(_ctx(), web.FetchUrlInput(
url="https://example.com/spa", render=True))
assert ei.value.code == "render_unavailable"
class TestMarkdownExtraction:
def test_html_to_text_reused(self):
text = webrender._html_to_text("<html><body><p>hello</p><script>x()</script></body></html>")
assert "hello" in text
assert "x()" not in text