fix(security): consolidation & securite phase 1 (#84, BUG-021 a BUG-034)
CI / lint (push) Successful in 1m20s
CI / security (push) Successful in 47s
CI / test (push) Successful in 2m21s
CI / build (push) Successful in 43s
CI / e2e (push) Successful in 10m48s

- sanitizer XSS serveur (markdown + page de partage) [BUG-021/022]
- rate-limit/lockout MFA [BUG-023]
- isolation vaults par segments [BUG-024]
- caps regex ReDoS [BUG-025]
- SSRF webhooks + secrets externalises [BUG-026]
- rotation/revocation des jetons [BUG-027]
- politique de mot de passe + invalidation sessions [BUG-028]
- verrous users.json [BUG-029]
- IP reelle dans les audits [BUG-030]
- rate-limit par compte [BUG-031]
- symlinks hors vault ignores [BUG-032]
- recherche simple via inverted index [BUG-033]
- token en memoire + cookie HttpOnly, CSP durcie [BUG-034]

Tests: pytest 961 passed / 6 skipped, ruff 0, mypy 0, frontend vert.
This commit is contained in:
2026-09-13 10:51:42 -04:00
parent d47fcbf2be
commit 162a5b4acc
29 changed files with 1690 additions and 281 deletions
+9
View File
@@ -16,8 +16,17 @@ OBSIGATE_ADMIN_PASSWORD=chab30
# Rate limiting
# OBSIGATE_LOGIN_MAX_ATTEMPTS=10
# OBSIGATE_ACCOUNT_MAX_ATTEMPTS=10
# OBSIGATE_LOGIN_WINDOW_SECONDS=900
# IP client derrière un reverse proxy (fait confiance à X-Forwarded-For)
# OBSIGATE_TRUST_PROXY=false
# Webhooks : sécurité SSRF
# OBSIGATE_WEBHOOK_ALLOW_HTTP=false # autoriser http:// (défaut : HTTPS requis)
# OBSIGATE_WEBHOOK_ALLOW_PRIVATE=false # autoriser les IP privées/boucle
# Secret d'un webhook : OBSIGATE_WEBHOOK_SECRET_<ID_WEBHOOK_EN_MAJUSCULES>
# Watcher
# OBSIGATE_WATCHER_ENABLED=true
# OBSIGATE_WATCHER_USE_POLLING=false
+40
View File
@@ -12,6 +12,46 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
### Sécurité
- **#84 Consolidation & sécurité — phase 1 (BUG-021 → BUG-034)** — traitement des
vulnérabilités de la revue statique du 2026-09-13 :
- **BUG-021/022 — XSS stocké** : nouveau sanitizer serveur en liste blanche
(`backend/services/sanitizer.py`, stdlib) appliqué au rendu markdown ; échappement
systématique du `title`, du frontmatter et du JSON de la page publique `/s/{token}`
(`</script>` neutralisé). Tests : `tests/test_security_hardening.py` (9).
- **BUG-023 — brute-force MFA** : rate-limit IP + compte et verrouillage de compte sur
`mfa/totp/verify`, `mfa/recovery` et `mfa/webauthn/verify` (`_enforce_mfa_rate_limit`).
- **BUG-024 — traversal inter-vaults** : `resolve_safe_path` compare désormais les chemins
par **segment** (`Path.relative_to` + repli casse-insensible), plus par préfixe de chaîne :
`vault` ne peut plus lire `vault-evil`.
- **BUG-025 — ReDoS** : validation des regex utilisateur (longueur max, rejet des
quantificateurs imbriqués/backrefs), contenu tronqué et nombre de matchs plafonné
(`backend/services/regex_safety.py`), appliqué à la recherche avancée et au find/replace.
- **BUG-026 — SSRF webhooks** : validation d'URL (HTTPS par défaut, IP privées/boucle
interdites, résolution DNS vérifiée au dispatch, redirections non suivies) et
**externalisation du secret** dans `data/webhook_secrets.json` (0600) ou variable
`OBSIGATE_WEBHOOK_SECRET_<ID>` (plus de secret en clair dans `webhooks.json`).
- **BUG-027 — sessions** : rotation du refresh token à chaque usage, révocation du JTI de
l'access token au logout, et vérification de la révocation dans le middleware.
- **BUG-028 — politique de mot de passe** : validation centralisée (8–128 caractères) à la
création, à la modification admin et au changement ; `password_changed_at` invalide tous
les jetons émis avant un changement de mot de passe.
- **BUG-029 — race `users.json`** : verrou `threading.RLock` autour des cycles
lecture-modification-écriture.
- **BUG-030 — audits** : l'adresse IP réelle du client (`X-Forwarded-For` si
`OBSIGATE_TRUST_PROXY=true`) est injectée dans `current_user` et consignée dans les audits.
- **BUG-031 — rate-limiter** : budget **par compte** en plus du budget par IP (rotation d'IP
neutralisée) ; limite mono-process documentée.
- **BUG-032 — indexation** : `_scan_vault` utilise `os.walk(followlinks=False)` et refuse
tout symlink sortant de la racine du vault.
- **BUG-033 — recherche O(N)** : la recherche simple et l'outil IA `search_fulltext`
utilisent l'inverted index (repli sur le scan pendant la construction).
- **BUG-034 — CSP & jetons** : le token d'accès n'est plus persisté dans `sessionStorage`
(mémoire + cookie `HttpOnly`) ; directives CSP durcies (`object-src`, `base-uri`,
`form-action`, `frame-ancestors`). *Reste : migration CSP par nonce (exige la conversion
des gestionnaires d'événements inline).*
### Ajouté
- **#82 Assistant IA — menu `@` instantané** — nouvel endpoint `GET /api/vault/{vault}/paths`
+4
View File
@@ -275,7 +275,11 @@ Un compte **admin** connecté voit une icône 🛡️ dans le header : liste, cr
| `OBSIGATE_ACCESS_TOKEN_TTL` | Durée de vie token JWT (secondes) | `3600` |
| `OBSIGATE_REFRESH_TOKEN_TTL` | Durée de vie refresh token (secondes) | `2592000` |
| `OBSIGATE_LOGIN_MAX_ATTEMPTS` | Tentatives de login max par IP | `10` |
| `OBSIGATE_ACCOUNT_MAX_ATTEMPTS` | Tentatives de login max par compte | `10` |
| `OBSIGATE_LOGIN_WINDOW_SECONDS` | Fenêtre de rate limiting (secondes) | `900` |
| `OBSIGATE_TRUST_PROXY` | Faire confiance à `X-Forwarded-For` pour l'IP client (reverse proxy) | `false` |
| `OBSIGATE_WEBHOOK_ALLOW_HTTP` | Autoriser les webhooks non HTTPS | `false` |
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Autoriser les webhooks vers des adresses privées/boucle | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Taille max des PDF extraits (text indexation) | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | Timeout extraction PDF (secondes) | `30` |
+4
View File
@@ -313,7 +313,11 @@ When an **admin** account is logged in, a 🛡️ icon appears in the header. Cl
| `OBSIGATE_ACCESS_TOKEN_TTL` | JWT token lifetime (seconds) | `3600` |
| `OBSIGATE_REFRESH_TOKEN_TTL` | Refresh token lifetime (seconds) | `2592000` |
| `OBSIGATE_LOGIN_MAX_ATTEMPTS` | Max login attempts per IP | `10` |
| `OBSIGATE_ACCOUNT_MAX_ATTEMPTS` | Max login attempts per account | `10` |
| `OBSIGATE_LOGIN_WINDOW_SECONDS` | Rate limiting window (seconds) | `900` |
| `OBSIGATE_TRUST_PROXY` | Trust `X-Forwarded-For` for the client IP (reverse proxy) | `false` |
| `OBSIGATE_WEBHOOK_ALLOW_HTTP` | Allow non-HTTPS webhook targets | `false` |
| `OBSIGATE_WEBHOOK_ALLOW_PRIVATE` | Allow webhooks to private/loopback addresses | `false` |
| `OBSIGATE_PDF_MAX_SIZE_MB` | Max PDF size for text extraction | `50` |
| `OBSIGATE_PDF_EXTRACT_TIMEOUT` | PDF extraction timeout (seconds) | `30` |
+8 -3
View File
@@ -62,16 +62,21 @@ def create_access_token(user: dict) -> str:
return jwt.encode(payload, get_secret_key(), algorithm=ALGORITHM)
def create_refresh_token(username: str) -> tuple:
"""Create a JWT refresh token. Returns (token_string, jti)."""
def create_refresh_token(username: str, remember: bool = False) -> tuple:
"""Create a JWT refresh token. Returns (token_string, jti).
``remember`` is carried as a claim so token rotation can preserve the
30-day vs 7-day lifetime chosen at login.
"""
now = int(time.time())
jti = str(uuid.uuid4())
payload = {
"sub": username,
"jti": jti,
"iat": now,
"exp": now + REFRESH_TOKEN_EXPIRE_SECONDS,
"exp": now + (2592000 if remember else REFRESH_TOKEN_EXPIRE_SECONDS),
"type": "refresh",
"remember": remember,
}
return jwt.encode(payload, get_secret_key(), algorithm=ALGORITHM), jti
+21 -1
View File
@@ -8,7 +8,9 @@ import os
from fastapi import Depends, HTTPException, Request
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
from .jwt_handler import decode_token
from backend.services.net import get_client_ip
from .jwt_handler import decode_token, is_token_revoked
from .user_store import get_user
logger = logging.getLogger("obsigate.auth.middleware")
@@ -42,6 +44,7 @@ def get_current_user(
"vaults": ["*"],
"active": True,
"_token_vaults": ["*"],
"_request_ip": get_client_ip(request),
}
token = None
@@ -57,14 +60,31 @@ def get_current_user(
if not payload or payload.get("type") != "access":
return None
# BUG-027: access tokens revoked at logout must be rejected immediately.
jti = payload.get("jti")
if jti and is_token_revoked(jti):
return None
user = get_user(payload["sub"])
if not user or not user.get("active"):
return None
# BUG-028: a password change invalidates every token issued before it.
pca = user.get("password_changed_at")
iat = payload.get("iat")
if pca is not None and iat is not None:
try:
if int(iat) < int(float(pca)):
return None
except (TypeError, ValueError):
return None
# Attach vault permissions from the token (snapshot at login time)
user["_token_vaults"] = payload.get("vaults", [])
# Attach the token id for per-token rate limiting (AI tool layer).
user["_token_jti"] = payload.get("jti")
# BUG-030: expose the real client IP to the audit log.
user["_request_ip"] = get_client_ip(request)
return user
+26
View File
@@ -13,6 +13,32 @@ ph = PasswordHasher(
salt_len=16,
)
# Password policy (BUG-028). Applied by the API validators at account creation
# and password change so the rules stay consistent across both paths.
MIN_PASSWORD_LENGTH = 8
MAX_PASSWORD_LENGTH = 128
def validate_password_strength(password: str) -> str:
"""Validate a plaintext password against the project policy.
Args:
password: Candidate password.
Returns:
The password unchanged when valid.
Raises:
ValueError: When the password is too short, too long or blank.
"""
if password is None or len(password) < MIN_PASSWORD_LENGTH:
raise ValueError(f"Minimum {MIN_PASSWORD_LENGTH} caractères")
if len(password) > MAX_PASSWORD_LENGTH:
raise ValueError(f"Maximum {MAX_PASSWORD_LENGTH} caractères")
if not password.strip():
raise ValueError("Le mot de passe ne peut pas être vide")
return password
def hash_password(password: str) -> str:
"""Hash a password with Argon2id."""
+116 -16
View File
@@ -8,9 +8,12 @@ import re
from fastapi import APIRouter, Body, Depends, HTTPException, Request, Response
from pydantic import BaseModel, validator
from backend.ratelimit import is_rate_limited
from backend.ratelimit import is_account_rate_limited, is_rate_limited
from backend.ratelimit import record_account_failure as rl_record_account_failure
from backend.ratelimit import record_account_success as rl_record_account_success
from backend.ratelimit import record_failure as rl_record_failure
from backend.ratelimit import record_success as rl_record_success
from backend.services.net import get_client_ip
from .jwt_handler import (
ACCESS_TOKEN_EXPIRE_SECONDS,
@@ -29,7 +32,7 @@ from .mfa import (
verify_totp,
)
from .middleware import is_auth_enabled, require_admin, require_auth
from .password import hash_password, verify_password
from .password import hash_password, validate_password_strength, verify_password
from .user_store import (
create_user,
delete_user,
@@ -61,9 +64,7 @@ class ChangePasswordRequest(BaseModel):
@validator("new_password")
def password_strength(cls, v):
if len(v) < 8:
raise ValueError("Minimum 8 caractères")
return v
return validate_password_strength(v)
class CreateUserRequest(BaseModel):
@@ -73,6 +74,10 @@ class CreateUserRequest(BaseModel):
role: str = "user"
vaults: list[str] = []
@validator("password")
def password_valid(cls, v):
return validate_password_strength(v)
@validator("username")
def username_valid(cls, v):
if not re.match(r"^[a-zA-Z0-9_-]{2,32}$", v):
@@ -93,6 +98,12 @@ class UpdateUserRequest(BaseModel):
password: str | None = None
role: str | None = None
@validator("password")
def password_valid(cls, v):
if v is None:
return v
return validate_password_strength(v)
# ── Public endpoints ──────────────────────────────────────────────────
@@ -128,16 +139,21 @@ async def login(body: LoginRequest, response: Response, request: Request):
raise HTTPException(403, "Compte désactivé")
# IP-based rate limiting (10 failures / 15 min per IP)
client_ip = request.client.host if request.client else "unknown"
client_ip = get_client_ip(request)
if is_rate_limited(client_ip):
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
# BUG-031: per-account budget still applies when the attacker rotates IPs.
if is_account_rate_limited(body.username):
raise HTTPException(429, "Trop de tentatives sur ce compte (15min)")
if is_locked(body.username):
raise HTTPException(429, "Compte temporairement verrouillé (15min)")
if not verify_password(body.password, user["password_hash"]):
attempts = record_login_failure(body.username)
rl_attempts, rl_remaining = rl_record_failure(client_ip)
rl_record_account_failure(body.username)
remaining = max(0, 5 - attempts)
detail = "Identifiants invalides"
if 0 < remaining <= 2:
@@ -164,9 +180,10 @@ async def login(body: LoginRequest, response: Response, request: Request):
def _issue_tokens(user: dict, username: str, remember_me: bool, response: Response) -> dict:
"""Issue JWT tokens after successful authentication (password or MFA verified)."""
record_login_success(username)
rl_record_account_success(username)
access_token = create_access_token(user)
refresh_token, refresh_jti = create_refresh_token(username)
refresh_token, refresh_jti = create_refresh_token(username, remember=remember_me)
import os
max_age = 2592000 if remember_me else 604800 # 30d or 7d
@@ -208,6 +225,8 @@ async def refresh_token_endpoint(request: Request, response: Response):
"""Renew access token via refresh token cookie.
Called automatically by the frontend when the access token expires.
The refresh token is rotated on every use (BUG-027) and rejected if it
predates the user's last password change (BUG-028).
"""
refresh_tok = request.cookies.get("refresh_token")
if not refresh_tok:
@@ -224,11 +243,38 @@ async def refresh_token_endpoint(request: Request, response: Response):
if not user or not user.get("active"):
raise HTTPException(401, "Utilisateur introuvable ou inactif")
# BUG-028: reject refresh tokens issued before the last password change.
pca = user.get("password_changed_at")
iat = payload.get("iat")
if pca is not None and iat is not None:
try:
stale = int(iat) < int(float(pca))
except (TypeError, ValueError):
stale = True
if stale:
raise HTTPException(401, "Session expirée, veuillez vous reconnecter")
import os
secure = os.environ.get("OBSIGATE_SECURE_COOKIES", "false").lower() == "true"
remember_me = bool(payload.get("remember", False))
# BUG-027: rotate the refresh token — the old one is now single-use.
revoke_token(payload["jti"])
new_refresh_token, _new_jti = create_refresh_token(user["username"], remember=remember_me)
max_age = 2592000 if remember_me else 604800
response.set_cookie(
key="refresh_token",
value=new_refresh_token,
max_age=max_age,
httponly=True,
samesite="strict",
secure=secure,
path="/api/auth/refresh",
)
new_access_token = create_access_token(user)
# Update cookies
import os
secure = os.environ.get("OBSIGATE_SECURE_COOKIES", "false").lower() == "true"
response.set_cookie(
key="access_token",
value=new_access_token,
@@ -251,7 +297,7 @@ async def logout(
request: Request,
response: Response,
):
"""Logout: revoke refresh token and delete cookies."""
"""Logout: revoke refresh and access tokens, then delete cookies."""
refresh_tok = request.cookies.get("refresh_token")
if refresh_tok:
payload = decode_token(refresh_tok)
@@ -261,6 +307,21 @@ async def logout(
except Exception:
pass # token already revoked
# BUG-027: revoke the access token too, otherwise it stays valid until expiry.
access_tok = None
auth_header = request.headers.get("authorization", "")
if auth_header.lower().startswith("bearer "):
access_tok = auth_header[7:].strip()
if not access_tok:
access_tok = request.cookies.get("access_token")
if access_tok:
access_payload = decode_token(access_tok)
if access_payload and access_payload.get("type") == "access":
try:
revoke_token(access_payload["jti"])
except Exception:
pass
response.delete_cookie("refresh_token", path="/api/auth/refresh")
response.delete_cookie("access_token", path="/")
response.delete_cookie("access_token", path="/api") # just in case
@@ -313,19 +374,51 @@ async def patch_me(req: UpdateMeRequest, current_user=Depends(require_auth)):
@router.post("/change-password")
async def change_password(
req: ChangePasswordRequest,
response: Response,
current_user=Depends(require_auth),
):
"""Change own password."""
"""Change own password.
BUG-028: changing the password invalidates all previously issued tokens;
a fresh pair is issued to keep the current session alive.
"""
user = get_user(current_user["username"])
assert user is not None, f"User {current_user['username']} not found"
if not verify_password(req.current_password, user["password_hash"]):
raise HTTPException(400, "Mot de passe actuel incorrect")
update_user(current_user["username"], {"password": req.new_password})
return {"message": "Mot de passe mis à jour"}
updated = get_user(current_user["username"])
result: dict = {"message": "Mot de passe mis à jour"}
if updated is not None:
result.update(_issue_tokens(updated, updated["username"], False, response))
return result
# ── MFA endpoints ────────────────────────────────────────────────────
def _enforce_mfa_rate_limit(request: Request, username: str) -> str:
"""Reject MFA attempts from a rate-limited IP or on a locked account.
BUG-023: the second-factor endpoints were previously unprotected, making
the 6-digit TOTP brute-forceable. Returns the resolved client IP.
"""
client_ip = get_client_ip(request)
if is_rate_limited(client_ip):
raise HTTPException(429, "Trop de tentatives depuis cette adresse IP (15min)")
if is_account_rate_limited(username):
raise HTTPException(429, "Trop de tentatives sur ce compte (15min)")
if is_locked(username):
raise HTTPException(429, "Compte temporairement verrouillé (15min)")
return client_ip
def _record_mfa_failure(client_ip: str, username: str) -> None:
"""Record a failed MFA attempt for the IP, the account and the lockout."""
record_login_failure(username)
rl_record_failure(client_ip)
rl_record_account_failure(username)
class MfaVerifyRequest(BaseModel):
username: str
code: str
@@ -601,6 +694,8 @@ async def mfa_webauthn_verify(
from .user_store import get_user, update_user
from .webauthn_mfa import complete_authentication
client_ip = _enforce_mfa_rate_limit(request, body.username)
user = get_user(body.username)
if not user:
hash_password("dummy_timing_protection")
@@ -616,8 +711,10 @@ async def mfa_webauthn_verify(
raise ValueError("Credential non enregistré")
new_count = complete_authentication(body.username, body.credential, stored)
except ValueError as e:
_record_mfa_failure(client_ip, body.username)
raise HTTPException(401, str(e))
except Exception as e:
_record_mfa_failure(client_ip, body.username)
logger.warning(f"WebAuthn verification failed for {body.username}: {e}")
raise HTTPException(401, "Vérification WebAuthn échouée")
@@ -627,7 +724,6 @@ async def mfa_webauthn_verify(
c["sign_count"] = new_count
update_user(body.username, {"webauthn_credentials": updated})
client_ip = request.client.host if request.client else "unknown"
rl_record_success(client_ip)
logger.info(f"User '{body.username}' logged in via WebAuthn")
return _issue_tokens(user, body.username, body.remember_me, response)
@@ -655,6 +751,8 @@ async def mfa_totp_verify(body: MfaVerifyRequest, response: Response, request: R
"""
from .user_store import get_user
client_ip = _enforce_mfa_rate_limit(request, body.username)
user = get_user(body.username)
if not user:
# Timing-safe: simulate work
@@ -665,10 +763,10 @@ async def mfa_totp_verify(body: MfaVerifyRequest, response: Response, request: R
raise HTTPException(400, "MFA non activé pour cet utilisateur")
if not verify_totp(user["mfa_secret"], body.code):
_record_mfa_failure(client_ip, body.username)
raise HTTPException(401, "Code TOTP invalide")
# Clear IP rate limit on success
client_ip = request.client.host if request.client else "unknown"
rl_record_success(client_ip)
return _issue_tokens(user, body.username, body.remember_me, response)
@@ -682,6 +780,8 @@ async def mfa_recovery_login(body: MfaRecoveryRequest, response: Response, reque
"""
from .user_store import get_user, update_user
client_ip = _enforce_mfa_rate_limit(request, body.username)
user = get_user(body.username)
if not user:
hash_password("dummy_timing_protection")
@@ -696,6 +796,7 @@ async def mfa_recovery_login(body: MfaRecoveryRequest, response: Response, reque
idx = verify_recovery_code(body.recovery_code, hashed_codes)
if idx is None:
_record_mfa_failure(client_ip, body.username)
raise HTTPException(401, "Code de récupération invalide")
# Remove used recovery code (single-use)
@@ -703,7 +804,6 @@ async def mfa_recovery_login(body: MfaRecoveryRequest, response: Response, reque
update_user(body.username, {"mfa_recovery_codes": hashed_codes})
# Clear IP rate limit
client_ip = request.client.host if request.client else "unknown"
rl_record_success(client_ip)
logger.info(f"User '{body.username}' logged in via recovery code")
+64 -49
View File
@@ -6,6 +6,7 @@
import json
import logging
import shutil
import threading
import uuid
from datetime import datetime, timedelta, timezone
from pathlib import Path
@@ -16,6 +17,12 @@ logger = logging.getLogger("obsigate.auth.users")
USERS_FILE = Path("data/users.json")
# Serialises read-modify-write cycles on users.json. ``RLock`` because a few
# helpers (e.g. ``record_login_failure``) call other mutators while holding it.
# BUG-029: without this, concurrent MFA enable + password change could lose one
# of the two updates (last writer wins).
_users_lock = threading.RLock()
def _read() -> dict:
"""Read users.json. Returns empty structure if file doesn't exist."""
@@ -75,26 +82,28 @@ def create_user(
display_name: str | None = None,
) -> dict:
"""Create a new user. Raises ValueError if username already taken."""
data = _read()
if username in data["users"]:
raise ValueError(f"User '{username}' already exists")
with _users_lock:
data = _read()
if username in data["users"]:
raise ValueError(f"User '{username}' already exists")
user = {
"id": str(uuid.uuid4()),
"username": username,
"display_name": display_name or username,
"password_hash": hash_password(password),
"role": role,
"vaults": vaults or [],
"active": True,
"language": "fr", # default UI language
"created_at": datetime.now(timezone.utc).isoformat(),
"last_login": None,
"failed_attempts": 0,
"locked_until": None,
}
data["users"][username] = user
_write(data)
user = {
"id": str(uuid.uuid4()),
"username": username,
"display_name": display_name or username,
"password_hash": hash_password(password),
"role": role,
"vaults": vaults or [],
"active": True,
"language": "fr", # default UI language
"created_at": datetime.now(timezone.utc).isoformat(),
"password_changed_at": datetime.now(timezone.utc).timestamp(),
"last_login": None,
"failed_attempts": 0,
"locked_until": None,
}
data["users"][username] = user
_write(data)
logger.info(f"Created user '{username}' (role={role})")
return {k: v for k, v in user.items() if k != "password_hash"}
@@ -105,28 +114,33 @@ def update_user(username: str, updates: dict) -> dict:
Forbidden fields (id, username, created_at) are silently ignored.
If 'password' is in updates, it's hashed and stored as password_hash.
"""
data = _read()
if username not in data["users"]:
raise ValueError(f"User '{username}' not found")
with _users_lock:
data = _read()
if username not in data["users"]:
raise ValueError(f"User '{username}' not found")
forbidden = {"id", "username", "created_at"}
safe_updates = {k: v for k, v in updates.items() if k not in forbidden}
forbidden = {"id", "username", "created_at"}
safe_updates = {k: v for k, v in updates.items() if k not in forbidden}
if "password" in safe_updates:
safe_updates["password_hash"] = hash_password(safe_updates.pop("password"))
if "password" in safe_updates:
safe_updates["password_hash"] = hash_password(safe_updates.pop("password"))
# BUG-028: invalidate every token issued before this change.
safe_updates["password_changed_at"] = datetime.now(timezone.utc).timestamp()
data["users"][username].update(safe_updates)
_write(data)
return {k: v for k, v in data["users"][username].items() if k != "password_hash"}
data["users"][username].update(safe_updates)
_write(data)
result = {k: v for k, v in data["users"][username].items() if k != "password_hash"}
return result
def delete_user(username: str):
"""Delete a user. Raises ValueError if not found."""
data = _read()
if username not in data["users"]:
raise ValueError(f"User '{username}' not found")
del data["users"][username]
_write(data)
with _users_lock:
data = _read()
if username not in data["users"]:
raise ValueError(f"User '{username}' not found")
del data["users"][username]
_write(data)
logger.info(f"Deleted user '{username}'")
@@ -144,24 +158,25 @@ def record_login_failure(username: str) -> int:
After 5 failures, locks the account for 15 minutes.
"""
data = _read()
user = data["users"].get(username)
if not user:
return 0
with _users_lock:
data = _read()
user = data["users"].get(username)
if not user:
return 0
attempts = user.get("failed_attempts", 0) + 1
updates = {"failed_attempts": attempts}
attempts = user.get("failed_attempts", 0) + 1
updates = {"failed_attempts": attempts}
# Lock after 5 failed attempts (15 minutes)
if attempts >= 5:
locked_until = (
datetime.now(timezone.utc) + timedelta(minutes=15)
).isoformat()
updates["locked_until"] = locked_until
logger.warning(f"Account '{username}' locked after {attempts} failed attempts")
# Lock after 5 failed attempts (15 minutes)
if attempts >= 5:
locked_until = (
datetime.now(timezone.utc) + timedelta(minutes=15)
).isoformat()
updates["locked_until"] = locked_until
logger.warning(f"Account '{username}' locked after {attempts} failed attempts")
update_user(username, updates)
return attempts
update_user(username, updates)
return attempts
def is_locked(username: str) -> bool:
+95 -76
View File
@@ -423,94 +423,113 @@ def _scan_vault(vault_name: str, vault_path: str, vault_cfg: dict[str, Any] | No
logger.warning(f"Vault path does not exist: {vault_path}")
return {"files": [], "tags": {}, "path": vault_path, "paths": []}
for fpath in vault_root.rglob("*"):
# Skip ignored directories
if any(part in IGNORED_DIRS for part in fpath.relative_to(vault_root).parts):
continue
root_resolved = vault_root.resolve(strict=False)
rel_path_str = str(fpath.relative_to(vault_root)).replace("\\", "/")
# Add all paths (files and directories) to path index
if fpath.is_dir():
# BUG-032: walk without following symlinks and refuse any symlink that
# escapes the vault root, so external data can never be indexed/exposed.
for dirpath, dirnames, filenames in os.walk(vault_root, followlinks=False):
current_dir = Path(dirpath)
# Prune ignored and symlinked directories in place (no recursion).
dirnames[:] = [
d for d in dirnames
if d not in IGNORED_DIRS and not (current_dir / d).is_symlink()
]
for d in dirnames:
dpath = current_dir / d
rel_path_str = str(dpath.relative_to(vault_root)).replace("\\", "/")
paths.append({
"path": rel_path_str,
"name": fpath.name,
"name": d,
"type": "directory"
})
continue
# Files only from here
if not fpath.is_file():
continue
ext = fpath.suffix.lower()
# Also match extensionless files named like Dockerfile, Makefile
basename_lower = fpath.name.lower()
if ext not in SUPPORTED_EXTENSIONS and basename_lower not in ("dockerfile", "makefile", "cmakelists.txt"):
continue
# Add file to path index
paths.append({
"path": rel_path_str,
"name": fpath.name,
"type": "file"
})
try:
relative = fpath.relative_to(vault_root)
stat = fpath.stat()
modified = datetime.fromtimestamp(stat.st_mtime, tz=timezone.utc).isoformat()
# PDF handling — special path (binary, uses pdf_reader)
tags: list[str] = []
if ext == ".pdf":
from backend.pdf_reader import extract_pdf_metadata, extract_pdf_text
raw = extract_pdf_text(fpath, max_chars=100000)
pdf_meta = extract_pdf_metadata(fpath)
title = pdf_meta.get("title") or fpath.stem.replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
elif ext == ".excalidraw" or fpath.name.lower().endswith(".excalidraw.md"):
raw = fpath.read_text(encoding="utf-8", errors="replace")
raw = extract_excalidraw_indexable(raw)
title = fpath.stem.replace(".excalidraw", "").replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
else:
raw = fpath.read_text(encoding="utf-8", errors="replace")
title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
for fname in filenames:
fpath = current_dir / fname
if ext == ".md":
post = parse_markdown_file(raw)
tags = _extract_tags(post)
inline_tags = _extract_inline_tags(post.content)
tags = list(set(tags) | set(inline_tags))
title = _extract_title(post, fpath)
content_preview = post.content[:200].strip()
if fpath.is_symlink():
try:
target = fpath.resolve(strict=True)
except OSError:
continue
try:
target.relative_to(root_resolved)
except ValueError:
logger.warning(f"Skipping symlink outside vault: {fpath}")
continue
_extract_wikilinks_for_backlinks(
vault_name, str(relative).replace("\\", "/"),
title, post.content
)
rel_path_str = str(fpath.relative_to(vault_root)).replace("\\", "/")
files.append({
"path": str(relative).replace("\\", "/"),
"title": title,
"tags": tags,
"content_preview": content_preview,
"content": raw[:SEARCH_CONTENT_LIMIT],
"size": stat.st_size,
"modified": modified,
"extension": ext,
ext = fpath.suffix.lower()
# Also match extensionless files named like Dockerfile, Makefile
basename_lower = fpath.name.lower()
if ext not in SUPPORTED_EXTENSIONS and basename_lower not in ("dockerfile", "makefile", "cmakelists.txt"):
continue
# Add file to path index
paths.append({
"path": rel_path_str,
"name": fname,
"type": "file"
})
for tag in tags:
tag_counts[tag] = tag_counts.get(tag, 0) + 1
try:
relative = fpath.relative_to(vault_root)
stat = fpath.stat()
modified = datetime.fromtimestamp(stat.st_mtime, tz=timezone.utc).isoformat()
except PermissionError:
logger.debug(f"Permission denied, skipping {fpath}")
continue
except Exception as e:
logger.error(f"Error indexing {fpath}: {e}")
continue
# PDF handling — special path (binary, uses pdf_reader)
tags: list[str] = []
if ext == ".pdf":
from backend.pdf_reader import extract_pdf_metadata, extract_pdf_text
raw = extract_pdf_text(fpath, max_chars=100000)
pdf_meta = extract_pdf_metadata(fpath)
title = pdf_meta.get("title") or fpath.stem.replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
elif ext == ".excalidraw" or fpath.name.lower().endswith(".excalidraw.md"):
raw = fpath.read_text(encoding="utf-8", errors="replace")
raw = extract_excalidraw_indexable(raw)
title = fpath.stem.replace(".excalidraw", "").replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
else:
raw = fpath.read_text(encoding="utf-8", errors="replace")
title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = raw[:200].strip()
if ext == ".md":
post = parse_markdown_file(raw)
tags = _extract_tags(post)
inline_tags = _extract_inline_tags(post.content)
tags = list(set(tags) | set(inline_tags))
title = _extract_title(post, fpath)
content_preview = post.content[:200].strip()
_extract_wikilinks_for_backlinks(
vault_name, str(relative).replace("\\", "/"),
title, post.content
)
files.append({
"path": str(relative).replace("\\", "/"),
"title": title,
"tags": tags,
"content_preview": content_preview,
"content": raw[:SEARCH_CONTENT_LIMIT],
"size": stat.st_size,
"modified": modified,
"extension": ext,
})
for tag in tags:
tag_counts[tag] = tag_counts.get(tag, 0) + 1
except PermissionError:
logger.debug(f"Permission denied, skipping {fpath}")
continue
except Exception as e:
logger.error(f"Error indexing {fpath}: {e}")
continue
logger.info(f"Vault '{vault_name}': indexed {len(files)} files, {len(paths)} paths, {len(tag_counts)} unique tags")
return {"files": files, "tags": tag_counts, "path": vault_path, "paths": paths, "config": {}}
+40 -14
View File
@@ -136,6 +136,7 @@ from backend.services.mutations import (
restore_backup as service_restore_backup,
)
from backend.services.recent import humanize_mtime, list_recent
from backend.services.sanitizer import sanitize_html
from backend.services.search import advanced_search_vaults, list_paths, search_paths, search_vaults
from backend.services.search import list_tags as service_list_tags
from backend.services.vaults import (
@@ -678,7 +679,11 @@ class SecurityHeadersMiddleware(BaseHTTPMiddleware):
"connect-src 'self' blob: https://esm.sh https://unpkg.com https://cdnjs.cloudflare.com https://fonts.googleapis.com https://fonts.gstatic.com https://cdn.jsdelivr.net; "
"font-src 'self' data: https://fonts.gstatic.com https://esm.sh; "
"worker-src 'self' blob:; "
"frame-src 'self' blob:;"
"frame-src 'self' blob:; "
"object-src 'none'; "
"base-uri 'self'; "
"form-action 'self'; "
"frame-ancestors 'self';"
)
# Static assets are NOT content-hashed, so they must revalidate:
# ``immutable``/long max-age made Cloudflare and mobile browsers serve
@@ -1178,7 +1183,10 @@ def _render_markdown(raw_md: str, vault_name: str, current_file_path: Path | Non
# Add heading IDs for TOC navigation
rendered = _add_heading_ids(rendered)
# Sanitize: raw HTML in vault content must never reach the DOM (BUG-021).
rendered = sanitize_html(rendered)
return rendered
@@ -2657,14 +2665,15 @@ async def api_advanced_search(
Results include ``<mark>``-highlighted snippets and faceted tag/vault counts.
"""
loop = asyncio.get_event_loop()
return await loop.run_in_executor(
_search_executor,
partial(advanced_search_vaults, q, vault=vault, tag=tag,
search_fn = partial(advanced_search_vaults, q, vault=vault, tag=tag,
limit=limit, offset=offset, sort=sort,
case_sensitive=case_sensitive, whole_word=whole_word, regex=regex,
include_paths=include_paths, exclude_paths=exclude_paths,
created=created, modified=modified, size=size, semantic=semantic),
)
created=created, modified=modified, size=size, semantic=semantic)
try:
return await loop.run_in_executor(_search_executor, search_fn)
except ValueError as e:
raise HTTPException(400, str(e)) from e
@app.post("/api/search/replace", response_model=ReplaceResponse)
@@ -4012,9 +4021,23 @@ async def public_share_view(token: str):
title = post.metadata.get("title", file_path.stem)
# JSON-escape raw content for embedding in HTML
import json as _json
raw_json = _json.dumps(raw)
# Escape everything user-controlled before embedding in HTML/JS (BUG-022).
title_esc = html_mod.escape(str(title))
# Neutralise ``</script>`` in the JS string literal too.
title_download_js = (
_json.dumps(f"{title}.md")
.replace("<", "\\u003c")
.replace(">", "\\u003e")
.replace("&", "\\u0026")
)
# JSON-escape raw content for embedding in HTML, and neutralise ``</script>``.
raw_json = (
_json.dumps(raw)
.replace("<", "\\u003c")
.replace(">", "\\u003e")
.replace("&", "\\u0026")
)
fm_html = ""
if post.metadata:
fm_items = []
@@ -4028,12 +4051,15 @@ async def public_share_view(token: str):
v = "✓" if v else "✗"
elif v is None:
v = "—"
fm_items.append(f'<div class="fm-row"><span class="fm-key">{k}</span><span class="fm-val">{v}</span></div>')
fm_items.append(
f'<div class="fm-row"><span class="fm-key">{html_mod.escape(str(k))}</span>'
f'<span class="fm-val">{html_mod.escape(str(v))}</span></div>'
)
if fm_items:
fm_html = f'<div class="fm-section"><div class="fm-header">Frontmatter</div><div class="fm-body">{"".join(fm_items)}</div></div>'
return HTMLResponse(f"""<!DOCTYPE html><html lang="fr" data-theme="dark"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
<title>{title} — ObsiGate Share</title>
<title>{title_esc} — ObsiGate Share</title>
<style>
:root {{ --bg:#1a1a2e; --bg-card:#16213e; --text:#e0e0e0; --text-muted:#888; --accent:#6366f1; --border:#2a2a4a; --banner-bg:var(--accent); --banner-text:#fff; }}
[data-theme="light"] {{ --bg:#f8f9fa; --bg-card:#fff; --text:#1a1a2e; --text-muted:#666; --accent:#4f46e5; --border:#ddd; --banner-bg:#eef2ff; --banner-text:#4338ca; }}
@@ -4075,7 +4101,7 @@ body{{font-family:system-ui,-apple-system,sans-serif;background:var(--bg);color:
Document partagé via ObsiGate
</div>
<div class="toolbar">
<span class="toolbar-title">{title}</span>
<span class="toolbar-title">{title_esc}</span>
<button class="toolbar-btn" onclick="toggleTheme()" title="Thème clair/sombre">
<svg id="theme-icon-dark" xmlns="http://www.w3.org/2000/svg" width="15" height="15" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1 1 11.21 3 7 7 0 0 0 21 12.79z"/></svg>
<svg id="theme-icon-light" xmlns="http://www.w3.org/2000/svg" width="15" height="15" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" style="display:none"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg>
@@ -4094,7 +4120,7 @@ body{{font-family:system-ui,-apple-system,sans-serif;background:var(--bg);color:
<script>
function toggleTheme(){{var t=document.documentElement;var isDark=t.dataset.theme==="dark";t.dataset.theme=isDark?"light":"dark";document.getElementById("theme-icon-dark").style.display=isDark?"none":"";document.getElementById("theme-icon-light").style.display=isDark?"":"none";localStorage.setItem("obsigate-share-theme",t.dataset.theme)}}
(function(){{var s=localStorage.getItem("obsigate-share-theme");if(!s)s="dark";document.documentElement.dataset.theme=s;var isDark=s==="dark";document.getElementById("theme-icon-dark").style.display=isDark?"":"none";document.getElementById("theme-icon-light").style.display=isDark?"none":""}})();
function exportMD(){{var raw=JSON.parse(document.getElementById("raw-content").textContent);var b=new Blob([raw],{{type:"text/markdown"}});var a=document.createElement("a");a.href=URL.createObjectURL(b);a.download="{title}.md";a.click()}}
function exportMD(){{var raw=JSON.parse(document.getElementById("raw-content").textContent);var b=new Blob([raw],{{type:"text/markdown"}});var a=document.createElement("a");a.href=URL.createObjectURL(b);a.download={title_download_js};a.click()}}
</script></body></html>""")
+64 -15
View File
@@ -1,13 +1,21 @@
"""
IP-based rate limiter for authentication endpoints.
In-memory rate limiter for authentication endpoints.
Tracks failed login attempts per IP address with automatic
cleanup of expired entries. Complements the per-account lockout
in user_store.py.
Tracks failed attempts per IP **and** per account with automatic cleanup of
expired entries. The per-IP budget stops a single source; the per-account
budget (BUG-031) still throttles an attacker who rotates IPs. It complements
the per-account lockout in ``user_store.py``.
.. note::
The counters live in process memory only. They are **not** shared between
multiple workers/containers and are lost on restart. For a multi-node
deployment, front this service with a shared store (Redis) or a single
worker. This limitation is intentional and documented (BUG-031).
Configuration via environment variables:
OBSIGATE_LOGIN_MAX_ATTEMPTS Max failures per IP (default: 10)
OBSIGATE_LOGIN_WINDOW_SECONDS Lockout window in seconds (default: 900 = 15min)
OBSIGATE_LOGIN_MAX_ATTEMPTS Max failures per IP (default: 10)
OBSIGATE_ACCOUNT_MAX_ATTEMPTS Max failures per account (default: 10)
OBSIGATE_LOGIN_WINDOW_SECONDS Lockout window in seconds (default: 900)
"""
import logging
@@ -19,29 +27,37 @@ logger = logging.getLogger("obsigate.ratelimit")
# --- Configuration ---
MAX_ATTEMPTS = int(os.environ.get("OBSIGATE_LOGIN_MAX_ATTEMPTS", "10"))
ACCOUNT_MAX_ATTEMPTS = int(os.environ.get("OBSIGATE_ACCOUNT_MAX_ATTEMPTS", "10"))
WINDOW_SECONDS = int(os.environ.get("OBSIGATE_LOGIN_WINDOW_SECONDS", "900")) # 15 min
# --- In-memory store: {ip: [(timestamp, success_bool), ...]} ---
# --- In-memory stores: {key: [(timestamp, success_bool), ...]} ---
_ip_attempts: dict[str, list] = defaultdict(list)
_account_attempts: dict[str, list] = defaultdict(list)
_last_cleanup = time.time()
CLEANUP_INTERVAL = 60 # seconds
def _prune(store: dict[str, list], cutoff: float) -> None:
"""Drop expired entries from one store in place."""
expired = []
for key, attempts in store.items():
store[key] = [a for a in attempts if a[0] > cutoff]
if not store[key]:
expired.append(key)
for key in expired:
del store[key]
def _cleanup_expired():
"""Remove entries older than the window."""
"""Remove entries older than the window from both stores."""
global _last_cleanup
now = time.time()
if now - _last_cleanup < CLEANUP_INTERVAL:
return
_last_cleanup = now
cutoff = now - WINDOW_SECONDS
expired_ips = []
for ip, attempts in _ip_attempts.items():
_ip_attempts[ip] = [a for a in attempts if a[0] > cutoff]
if not _ip_attempts[ip]:
expired_ips.append(ip)
for ip in expired_ips:
del _ip_attempts[ip]
_prune(_ip_attempts, cutoff)
_prune(_account_attempts, cutoff)
def record_failure(ip: str) -> tuple[int, int]:
@@ -72,6 +88,37 @@ def is_rate_limited(ip: str) -> bool:
return failures >= MAX_ATTEMPTS
def record_account_failure(account: str) -> tuple[int, int]:
"""Record a failed attempt for an account, regardless of source IP.
Returns:
(current_failure_count, remaining_attempts)
"""
_cleanup_expired()
key = account.lower()
_account_attempts[key].append((time.time(), False))
failures = sum(1 for _, success in _account_attempts[key] if not success)
remaining = max(0, ACCOUNT_MAX_ATTEMPTS - failures)
if failures >= ACCOUNT_MAX_ATTEMPTS:
logger.warning(f"Account {account} rate-limited after {failures} failed attempts")
return failures, remaining
def record_account_success(account: str):
"""Clear the per-account rate limit state after a successful login."""
_cleanup_expired()
_account_attempts[account.lower()] = [(time.time(), True)]
def is_account_rate_limited(account: str) -> bool:
"""Check if an account has exceeded the per-account rate limit."""
_cleanup_expired()
failures = sum(
1 for _, success in _account_attempts.get(account.lower(), []) if not success
)
return failures >= ACCOUNT_MAX_ATTEMPTS
def get_status(ip: str | None = None) -> dict:
"""Get rate limit status for an IP (for diagnostics)."""
_cleanup_expired()
@@ -87,7 +134,9 @@ def get_status(ip: str | None = None) -> dict:
}
return {
"tracked_ips": len(_ip_attempts),
"tracked_accounts": len(_account_attempts),
"max_attempts": MAX_ATTEMPTS,
"account_max_attempts": ACCOUNT_MAX_ATTEMPTS,
"window_seconds": WINDOW_SECONDS,
"limited_ips": sum(
1 for ip_addr in _ip_attempts
+103 -50
View File
@@ -13,6 +13,11 @@ from sortedcontainers import SortedList
from backend import indexer as _indexer
from backend import semantic_search as _semantic
from backend.indexer import index
from backend.services.regex_safety import (
MAX_REGEX_MATCHES,
truncate_for_regex,
validate_regex,
)
logger = logging.getLogger("obsigate.search")
@@ -226,12 +231,16 @@ def _extract_regex_snippet(
if not content or not pattern_text:
return content[:200].strip() if content else ""
# BUG-025: bound the text scanned and the number of matches collected.
content = truncate_for_regex(content)
try:
validate_regex(pattern_text)
pattern = re.compile(pattern_text, re.IGNORECASE)
except re.error:
except (re.error, ValueError):
return _escape_html(content[:200].strip())
matches = list(pattern.finditer(content))
matches = list(pattern.finditer(content))[:MAX_REGEX_MATCHES]
if not matches:
return _escape_html(content[:200].strip())
@@ -729,60 +738,97 @@ def search(
query_lower = query.lower()
results: list[dict[str, Any]] = []
for vault_name, vault_data in index.items():
if vault_filter != "all" and vault_name != vault_filter:
inv = get_inverted_index()
use_index = (not inv.is_stale()) and inv.doc_count > 0
if use_index:
# BUG-033: retrieve candidates from the inverted index instead of
# scanning every document. Multi-term queries require all terms
# (a superset of exact-phrase matches), single terms use prefix
# expansion. Falls back to a full scan while the index is building.
if has_query:
terms = [t for t in tokenize(query) if t]
if not terms:
return []
doc_sets: list[set] = []
for term in terms:
term_docs: set = set(inv.word_index.get(term, {}).keys())
if len(term) >= MIN_PREFIX_LENGTH:
for expanded in inv.get_prefix_tokens(term):
term_docs.update(inv.word_index.get(expanded, {}).keys())
doc_sets.append(term_docs)
doc_keys = set.intersection(*doc_sets) if doc_sets else set()
else:
doc_keys = set(inv.doc_info.keys())
if vault_filter != "all":
doc_keys &= inv.vault_docs.get(vault_filter, set())
for tag in selected_tags:
doc_keys &= inv.tag_docs.get(tag.lower(), set())
candidates = [
(inv.doc_vault[dk], inv.doc_info[dk])
for dk in doc_keys
if dk in inv.doc_info
]
else:
candidates = [
(vault_name, file_info)
for vault_name, vault_data in index.items()
if vault_filter == "all" or vault_name == vault_filter
for file_info in vault_data["files"]
]
for vault_name, file_info in candidates:
# Tag filter: all selected tags must be present
if selected_tags and not all(tag in file_info["tags"] for tag in selected_tags):
continue
for file_info in vault_data["files"]:
# Tag filter: all selected tags must be present
if selected_tags and not all(tag in file_info["tags"] for tag in selected_tags):
continue
score = 0
snippet = file_info.get("content_preview", "")
score = 0
snippet = file_info.get("content_preview", "")
if has_query:
title_lower = file_info["title"].lower()
if has_query:
title_lower = file_info["title"].lower()
# Exact title match (highest weight)
if query_lower == title_lower:
score += 20
# Partial title match
elif query_lower in title_lower:
score += 10
# Exact title match (highest weight)
if query_lower == title_lower:
score += 20
# Partial title match
elif query_lower in title_lower:
score += 10
# Path match (folder/filename relevance)
if query_lower in file_info["path"].lower():
score += 5
# Path match (folder/filename relevance)
if query_lower in file_info["path"].lower():
score += 5
# Tag name match
for tag in file_info.get("tags", []):
if query_lower in tag.lower():
score += 3
break # count once per file
# Tag name match
for tag in file_info.get("tags", []):
if query_lower in tag.lower():
score += 3
break # count once per file
# Content match — use cached content (no disk I/O)
content = file_info.get("content", "")
content_lower = content.lower()
if query_lower in content_lower:
# Frequency-based scoring, capped to avoid over-weighting
occurrences = content_lower.count(query_lower)
score += min(occurrences, 10)
snippet = _extract_snippet(content, query)
else:
# Tag-only filter: all matching files get score 1
score = 1
# Content match — use cached content (no disk I/O)
content = file_info.get("content", "")
content_lower = content.lower()
if query_lower in content_lower:
# Frequency-based scoring, capped to avoid over-weighting
occurrences = content_lower.count(query_lower)
score += min(occurrences, 10)
snippet = _extract_snippet(content, query)
else:
# Tag-only filter: all matching files get score 1
score = 1
if score > 0:
results.append({
"vault": vault_name,
"path": file_info["path"],
"title": file_info["title"],
"tags": file_info["tags"],
"score": score,
"snippet": snippet,
"modified": file_info["modified"],
})
if score > 0:
results.append({
"vault": vault_name,
"path": file_info["path"],
"title": file_info["title"],
"tags": file_info["tags"],
"score": score,
"snippet": snippet,
"modified": file_info["modified"],
})
results.sort(key=lambda x: -x["score"])
return results[:limit]
@@ -908,12 +954,14 @@ def _passes_search_filters(
title = file_info.get("title", "")
content = file_info.get("content", "")
path = file_info.get("path", "")
search_text = f"{title} {content}"
# BUG-025: cap the text scanned by a user-supplied regex.
search_text = truncate_for_regex(f"{title} {content}")
search_text_norm = normalize_text(search_text)
# --- Regex mode ---
if regex and raw_query:
try:
validate_regex(raw_query)
flags = 0 if case_sensitive else re.IGNORECASE
if whole_word:
pattern = re.compile(rf"\b{raw_query}\b", flags)
@@ -921,7 +969,7 @@ def _passes_search_filters(
pattern = re.compile(raw_query, flags)
if not pattern.search(search_text):
return False
except re.error:
except (re.error, ValueError):
return False
return _passes_path_filters(path, include_paths, exclude_paths)
@@ -1138,6 +1186,11 @@ def advanced_search(
"""
t0 = time.monotonic()
query = query.strip() if query else ""
# BUG-025: reject oversized / catastrophic regex patterns up front.
if regex and query:
validate_regex(query)
parsed = _parse_advanced_query(query)
# Merge explicit tag_filter with parsed tag: operators
+23 -12
View File
@@ -433,22 +433,33 @@ def replace_in_files(
"""
import re as re_mod
from backend.services.regex_safety import MAX_REGEX_MATCHES, validate_regex
from backend.services.search import advanced_search_vaults
if not find:
raise ServiceError("Query is required", code="invalid_arguments", status=400)
search_results = advanced_search_vaults(
find,
vault=vault,
case_sensitive=case_sensitive,
whole_word=whole_word,
regex=regex,
include_paths=include_paths,
exclude_paths=exclude_paths,
limit=500,
sort="relevance",
)
# BUG-025: validate the pattern before it is compiled / applied in bulk.
if regex:
try:
validate_regex(find)
except ValueError as e:
raise ServiceError(str(e), code="invalid_arguments", status=400) from e
try:
search_results = advanced_search_vaults(
find,
vault=vault,
case_sensitive=case_sensitive,
whole_word=whole_word,
regex=regex,
include_paths=include_paths,
exclude_paths=exclude_paths,
limit=500,
sort="relevance",
)
except ValueError as e:
raise ServiceError(str(e), code="invalid_arguments", status=400) from e
if not search_results["results"]:
return {"matches": [], "total_matches": 0, "dry_run": dry_run}
@@ -480,7 +491,7 @@ def replace_in_files(
except OSError:
continue
occurrences = list(pattern.finditer(original))
occurrences = list(pattern.finditer(original))[:MAX_REGEX_MATCHES]
if not occurrences:
continue
+34
View File
@@ -0,0 +1,34 @@
"""Network helpers shared by the auth middleware and rate limiter."""
from __future__ import annotations
import os
from fastapi import Request
__all__ = ["get_client_ip", "is_trusted_proxy"]
def is_trusted_proxy() -> bool:
"""Whether ``X-Forwarded-For`` should be trusted (reverse proxy in front)."""
return os.environ.get("OBSIGATE_TRUST_PROXY", "false").lower() == "true"
def get_client_ip(request: Request) -> str:
"""Return the best-known client IP for *request*.
When ``OBSIGATE_TRUST_PROXY=true`` the left-most ``X-Forwarded-For`` entry
is used (the original client behind the proxy). Otherwise the socket peer
address is returned. BUG-030: this value feeds the audit log so attacks
remain traceable.
"""
if is_trusted_proxy():
forwarded = request.headers.get("x-forwarded-for")
if forwarded:
first = forwarded.split(",")[0].strip()
if first:
return first
real_ip = request.headers.get("x-real-ip")
if real_ip:
return real_ip.strip()
return request.client.host if request.client else "unknown"
+29 -12
View File
@@ -15,6 +15,27 @@ from backend.services.errors import ServiceError
logger = logging.getLogger("obsigate.services.paths")
def _is_within(resolved: Path, root: Path) -> bool:
"""Return True when *resolved* is *root* or lives below it.
The comparison is segment-aware so that a sibling directory whose name
merely shares a prefix (``vault`` vs ``vault-evil``) is rejected. A
case-insensitive fallback preserves the Windows / Docker behaviour where
the resolved casing can differ from the configured root.
"""
try:
resolved.relative_to(root)
return True
except ValueError:
pass
try:
resolved_parts = tuple(part.lower() for part in resolved.parts)
root_parts = tuple(part.lower() for part in root.parts)
except Exception:
return False
return resolved_parts[: len(root_parts)] == root_parts
def resolve_safe_path(vault_root: Path, relative_path: str | None) -> Path:
"""Resolve a vault-relative path, rejecting traversal outside the vault.
@@ -30,16 +51,12 @@ def resolve_safe_path(vault_root: Path, relative_path: str | None) -> Path:
logger.error(f"Path resolution error - vault_root: {vault_root}, relative_path: {relative_path}, error: {e}")
raise ServiceError(f"Path resolution error: {e!s}", code="path_error", status=500) from e
try:
resolved.relative_to(root)
except ValueError:
# Case-insensitive fallback for Windows / Docker path casing.
if not str(resolved).lower().startswith(str(root).lower()):
logger.warning(f"Path outside vault - vault: {root}, requested: {relative_path}, resolved: {resolved}")
raise ServiceError(
"Access denied: path outside vault",
code="path_outside_vault",
status=403,
details={"path": relative_path},
) from None
if not _is_within(resolved, root):
logger.warning(f"Path outside vault - vault: {root}, requested: {relative_path}, resolved: {resolved}")
raise ServiceError(
"Access denied: path outside vault",
code="path_outside_vault",
status=403,
details={"path": relative_path},
)
return resolved
+67
View File
@@ -0,0 +1,67 @@
"""Regex safety helpers (BUG-025).
User-supplied regular expressions are applied to large amounts of indexed
content. Python's :mod:`re` has no timeout, so a malicious pattern such as
``(a+)+$`` can pin a CPU for a long time (ReDoS). Without adding a native
dependency we mitigate by:
* bounding the pattern length,
* rejecting nested-quantifier constructs (the classic catastrophic form),
* capping the amount of text a single regex pass may scan,
* capping the number of matches collected.
"""
from __future__ import annotations
import re
__all__ = [
"MAX_PATTERN_LENGTH",
"MAX_REGEX_CONTENT",
"MAX_REGEX_MATCHES",
"truncate_for_regex",
"validate_regex",
]
MAX_PATTERN_LENGTH = 500
MAX_REGEX_CONTENT = 200_000
MAX_REGEX_MATCHES = 1000
# A quantified group whose body already contains a quantifier, immediately
# followed by another quantifier: ``(a+)+``, ``(.*)*``, ``(a+){2,}``, ...
_NESTED_QUANTIFIER_RE = re.compile(r"\([^()]*[+*][^()]*\)\s*(?:[+*?]|\{)")
# Backreferences combined with quantifiers are a common ReDoS vector too.
_BACKREF_QUANTIFIER_RE = re.compile(r"\\[1-9][0-9]*\s*(?:[+*]|\{)")
def validate_regex(pattern: str) -> str:
"""Validate a user-supplied regex against the safety policy.
Args:
pattern: Raw regex pattern.
Returns:
The pattern unchanged when acceptable.
Raises:
ValueError: When the pattern is empty, too long, or uses a construct
known to cause catastrophic backtracking.
"""
if not pattern:
raise ValueError("Expression régulière vide")
if len(pattern) > MAX_PATTERN_LENGTH:
raise ValueError(f"Expression régulière trop longue (max {MAX_PATTERN_LENGTH})")
if _NESTED_QUANTIFIER_RE.search(pattern) or _BACKREF_QUANTIFIER_RE.search(pattern):
raise ValueError("Expression régulière refusée (quantificateurs imbriqués)")
try:
re.compile(pattern)
except re.error as e:
raise ValueError(f"Expression régulière invalide : {e}") from e
return pattern
def truncate_for_regex(content: str, limit: int = MAX_REGEX_CONTENT) -> str:
"""Return at most *limit* characters to bound a single regex pass."""
if content and len(content) > limit:
return content[:limit]
return content
+286
View File
@@ -0,0 +1,286 @@
"""Whitelist HTML sanitizer used for untrusted markdown / AI output.
The markdown renderer runs with ``escape=False`` so raw HTML authored inside a
vault (or returned by a model) reaches the browser. This module scrubs the
rendered HTML against a strict whitelist of tags and attributes, drops
dangerous URL schemes and strips every event handler / ``style`` attribute.
Implemented with the standard library only (no third-party dependency) so the
runtime footprint stays unchanged. It is *not* a full HTML5 parser: it is a
conservative, allow-list based filter intended for already well-formed output
produced by mistune and the image/wikilink pre-processors.
"""
from __future__ import annotations
import html as _html
from html.parser import HTMLParser
__all__ = ["is_safe_url", "sanitize_html"]
# Tags whose *content* is discarded entirely (never rendered as text).
_DROP_CONTENT_TAGS = frozenset({
"script", "style", "iframe", "object", "embed", "template", "noscript",
"svg", "math", "applet", "form", "button", "select", "textarea", "option",
"frame", "frameset", "base", "link", "meta", "title", "head",
})
# Tags kept in the output (text content preserved for unknown tags).
_ALLOWED_TAGS = frozenset({
"a", "abbr", "b", "blockquote", "br", "caption", "code", "col", "colgroup",
"dd", "del", "details", "div", "dl", "dt", "em", "figcaption", "figure",
"h1", "h2", "h3", "h4", "h5", "h6", "hr", "i", "img", "input", "kbd", "li",
"mark", "ol", "p", "pre", "q", "s", "section", "small", "span", "strong",
"sub", "summary", "sup", "table", "tbody", "td", "tfoot", "th", "thead",
"time", "tr", "u", "ul", "video", "audio", "source", "track",
})
# Attributes allowed on any element.
_GLOBAL_ATTRS = frozenset({"class", "id", "title", "dir", "lang", "role"})
# Per-tag attribute whitelist (in addition to globals and ``data-*``).
_TAG_ATTRS: dict[str, frozenset[str]] = {
"a": frozenset({"href", "target", "rel", "name", "download"}),
"img": frozenset({"src", "alt", "width", "height", "loading"}),
"input": frozenset({"type", "checked", "disabled", "value"}),
"ol": frozenset({"start", "type", "reversed"}),
"ul": frozenset({"type"}),
"li": frozenset({"value"}),
"td": frozenset({"colspan", "rowspan", "align", "valign"}),
"th": frozenset({"colspan", "rowspan", "align", "valign", "scope"}),
"col": frozenset({"span", "width"}),
"colgroup": frozenset({"span"}),
"video": frozenset({"src", "controls", "width", "height", "loop", "muted",
"poster", "preload", "playsinline"}),
"audio": frozenset({"src", "controls", "loop", "muted", "preload"}),
"source": frozenset({"src", "type", "srcset", "media"}),
"track": frozenset({"src", "kind", "srclang", "label", "default"}),
"details": frozenset({"open"}),
"time": frozenset({"datetime"}),
"blockquote": frozenset({"cite"}),
"q": frozenset({"cite"}),
}
# URL-bearing attributes per tag, and whether ``data:`` URIs are acceptable.
_URL_ATTRS: dict[str, frozenset[str]] = {
"a": frozenset({"href"}),
"img": frozenset({"src"}),
"video": frozenset({"src", "poster"}),
"audio": frozenset({"src"}),
"source": frozenset({"src", "srcset"}),
"track": frozenset({"src"}),
"blockquote": frozenset({"cite"}),
"q": frozenset({"cite"}),
}
_SAFE_SCHEMES = frozenset({"http", "https", "mailto", "tel", "ftp"})
# Characters that browsers ignore inside a scheme (tab/newline/CR) and that
# could otherwise smuggle ``java\tscript:`` past a naive check.
_URL_STRIP_CHARS = "\t\n\r\x00"
def is_safe_url(value: str, *, allow_data: bool = False, tag: str = "") -> bool:
"""Return True when *value* is a URL with an allowed scheme.
Relative URLs (``/foo``, ``./foo``, ``#anchor``) are allowed. Dangerous
schemes such as ``javascript:`` and ``vbscript:`` are always rejected.
``data:`` URIs are only allowed for image/video/audio sources.
"""
if value is None:
return False
# Decode entities and strip whitespace/control chars before inspecting.
candidate = _html.unescape(str(value)).strip()
for ch in _URL_STRIP_CHARS:
candidate = candidate.replace(ch, "")
if not candidate:
return False
# Detect a scheme: ``scheme:`` where scheme is [a-zA-Z][a-zA-Z0-9+.-]*
lowered = candidate.lower()
if lowered.startswith("data:"):
if not allow_data:
return False
# Only media data URIs are permitted.
if tag in ("img",):
return lowered.startswith("data:image/")
if tag in ("video", "audio", "source", "track"):
return (
lowered.startswith("data:image/")
or lowered.startswith("data:video/")
or lowered.startswith("data:audio/")
)
return False
if lowered.startswith("blob:"):
return tag in ("img", "video", "audio", "source")
# No scheme at all (relative / fragment / protocol-relative) → safe.
colon = candidate.find(":")
slash = candidate.find("/")
if colon == -1 or (slash != -1 and slash < colon):
return True
# ``//host`` protocol-relative has no scheme.
if candidate.startswith("//"):
return True
scheme = lowered[:colon]
if not scheme or not scheme[0].isalpha():
return True # not a real scheme, treat as relative
return scheme in _SAFE_SCHEMES
class _Sanitizer(HTMLParser):
"""Rebuild HTML while dropping anything not explicitly allowed."""
def __init__(self) -> None:
super().__init__(convert_charrefs=True)
self._out: list[str] = []
# Stack of tag names currently open (only allowed tags).
self._open: list[str] = []
# Stack tracking dropped-content depth: each entry is the tag name.
self._suppress: list[str] = []
# -- helpers ----------------------------------------------------------
def _filter_attrs(self, tag: str, attrs: list[tuple[str, str | None]]) -> str:
allowed_extra = _TAG_ATTRS.get(tag, frozenset())
url_attrs = _URL_ATTRS.get(tag, frozenset())
parts: list[str] = []
seen: set[str] = set()
for name, value in attrs:
if value is None:
value = ""
lname = name.lower()
if lname in seen:
continue
seen.add(lname)
# Event handlers and style are never allowed.
if lname.startswith("on") or lname in ("style", "srcdoc", "formaction", "xlink:href"):
continue
if lname.startswith("data-") or lname.startswith("aria-"):
pass
elif lname not in _GLOBAL_ATTRS and lname not in allowed_extra:
continue
if lname in url_attrs:
allow_data = tag in ("img", "video", "audio", "source", "track")
# ``srcset`` may contain multiple comma-separated candidates.
if lname == "srcset":
if not _safe_srcset(value, tag):
continue
elif not is_safe_url(value, allow_data=allow_data, tag=tag):
continue
parts.append(f' {lname}="{_html.escape(value, quote=True)}"')
return "".join(parts)
def _emit_start(self, tag: str, attrs, self_closing: bool) -> None:
attrs_html = self._filter_attrs(tag, attrs)
if self_closing or tag in ("br", "hr", "img", "input", "col", "source", "track"):
self._out.append(f"<{tag}{attrs_html} />")
else:
self._out.append(f"<{tag}{attrs_html}>")
self._open.append(tag)
# -- HTMLParser callbacks --------------------------------------------
def handle_starttag(self, tag: str, attrs) -> None:
tag = tag.lower()
if tag in _DROP_CONTENT_TAGS:
self._suppress.append(tag)
return
if self._suppress:
return
if tag not in _ALLOWED_TAGS:
return # drop the tag, keep its text content
self._emit_start(tag, attrs, self_closing=False)
def handle_startendtag(self, tag: str, attrs) -> None:
tag = tag.lower()
if tag in _DROP_CONTENT_TAGS or self._suppress:
return
if tag not in _ALLOWED_TAGS:
return
self._emit_start(tag, attrs, self_closing=True)
def handle_endtag(self, tag: str) -> None:
tag = tag.lower()
if tag in _DROP_CONTENT_TAGS:
# Close the innermost matching suppress marker.
for i in range(len(self._suppress) - 1, -1, -1):
if self._suppress[i] == tag:
del self._suppress[i:]
break
return
if self._suppress:
return
if tag not in _ALLOWED_TAGS:
return
# Only close if currently open (tolerate malformed nesting).
if tag in self._open:
while self._open:
top = self._open.pop()
self._out.append(f"</{top}>")
if top == tag:
break
def handle_data(self, data: str) -> None:
if self._suppress:
return
self._out.append(_html.escape(data, quote=False))
def handle_comment(self, data: str) -> None:
return # comments are dropped
def handle_decl(self, decl: str) -> None:
return
def handle_pi(self, data: str) -> None:
return
def handle_entityref(self, name: str) -> None:
if self._suppress:
return
self._out.append(f"&{name};")
def handle_charref(self, name: str) -> None:
if self._suppress:
return
self._out.append(f"&#{name};")
def get_html(self) -> str:
# Close any tags left open by malformed input.
while self._open:
self._out.append(f"</{self._open.pop()}>")
return "".join(self._out)
def _safe_srcset(value: str, tag: str) -> bool:
"""Validate every candidate in a ``srcset`` attribute."""
for candidate in value.split(","):
candidate = candidate.strip()
if not candidate:
continue
url = candidate.split()[0] if candidate.split() else candidate
if not is_safe_url(url, allow_data=(tag == "img"), tag=tag):
return False
return True
def sanitize_html(html: str) -> str:
"""Return *html* with only whitelisted tags/attributes/schemes preserved.
Args:
html: Untrusted HTML (typically mistune output with ``escape=False``).
Returns:
Sanitized HTML string.
"""
if not html:
return ""
parser = _Sanitizer()
try:
parser.feed(html)
parser.close()
except Exception:
# Never let sanitization crash a request; fail closed to plain text.
return _html.escape(html)
return parser.get_html()
+172 -13
View File
@@ -4,6 +4,16 @@ Webhook management and dispatch for ObsiGate.
Webhooks are HTTP POST callbacks triggered on file/directory events.
Configuration is persisted in data/webhooks.json.
Security (BUG-026):
* target URLs are validated against SSRF (scheme + resolved IP must be public);
* HTTP is refused unless ``OBSIGATE_WEBHOOK_ALLOW_HTTP=true``;
* private/loopback/link-local targets are refused unless
``OBSIGATE_WEBHOOK_ALLOW_PRIVATE=true``;
* redirects are never followed;
* signing secrets are **not** stored in the public config file — they live in
``data/webhook_secrets.json`` (0600) or in an environment variable named
``OBSIGATE_WEBHOOK_SECRET_<ID>``.
Events: file_created, file_deleted, file_modified, file_renamed,
directory_created, directory_deleted, directory_renamed
"""
@@ -11,17 +21,22 @@ Events: file_created, file_deleted, file_modified, file_renamed,
import asyncio
import hashlib
import hmac
import ipaddress
import json
import logging
import os
import socket
import uuid
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
import aiohttp
logger = logging.getLogger("obsigate.webhooks")
WEBHOOKS_FILE = Path("data/webhooks.json")
WEBHOOK_SECRETS_FILE = Path("data/webhook_secrets.json")
VALID_EVENTS = {
"file_created", "file_deleted", "file_modified", "file_renamed",
@@ -29,6 +44,81 @@ VALID_EVENTS = {
}
def _allow_http() -> bool:
return os.environ.get("OBSIGATE_WEBHOOK_ALLOW_HTTP", "false").lower() == "true"
def _allow_private() -> bool:
return os.environ.get("OBSIGATE_WEBHOOK_ALLOW_PRIVATE", "false").lower() == "true"
def _is_public_ip(ip_str: str) -> bool:
"""True when *ip_str* is a globally routable unicast address."""
try:
ip = ipaddress.ip_address(ip_str)
except ValueError:
return False
return not (
ip.is_private or ip.is_loopback or ip.is_link_local or ip.is_reserved
or ip.is_multicast or ip.is_unspecified
)
def validate_webhook_url(url: str) -> str:
"""Validate the URL syntax/scheme and literal-IP safety at config time.
Raises:
ValueError: When the URL is malformed or points at an obviously
forbidden scheme/host.
"""
if not url or not isinstance(url, str):
raise ValueError("URL requise")
parsed = urlparse(url)
if parsed.scheme not in ("http", "https"):
raise ValueError("L'URL doit utiliser http ou https")
if parsed.scheme == "http" and not _allow_http():
raise ValueError("HTTPS requis (définir OBSIGATE_WEBHOOK_ALLOW_HTTP=true pour autoriser http)")
host = parsed.hostname
if not host:
raise ValueError("Hôte manquant dans l'URL")
# Reject literal private/loopback IPs immediately (no DNS needed).
try:
ip = ipaddress.ip_address(host)
except ValueError:
return url # hostname — resolved and checked at dispatch time
if not _allow_private() and not _is_public_ip(str(ip)):
raise ValueError("Adresse privée/interne refusée")
return url
def is_safe_target(url: str) -> bool:
"""Full SSRF check performed right before dispatch (resolves the host).
Returns False when the URL is malformed, the scheme is forbidden, or any
resolved address is private/loopback/reserved.
"""
try:
validate_webhook_url(url)
except ValueError:
return False
if _allow_private():
return True
parsed = urlparse(url)
host = parsed.hostname or ""
port = parsed.port or (443 if parsed.scheme == "https" else 80)
try:
infos = socket.getaddrinfo(host, port, proto=socket.IPPROTO_TCP)
except socket.gaierror:
logger.warning(f"Webhook target host could not be resolved: {host}")
return False
for info in infos:
addr = str(info[4][0])
if not _is_public_ip(addr):
logger.warning(f"Webhook target resolves to a non-public address ({addr}); blocked")
return False
return True
def _read() -> list:
if not WEBHOOKS_FILE.exists():
return []
@@ -45,35 +135,94 @@ def _write(webhooks: list):
tmp.replace(WEBHOOKS_FILE)
def _read_secrets() -> dict:
if not WEBHOOK_SECRETS_FILE.exists():
return {}
try:
return json.loads(WEBHOOK_SECRETS_FILE.read_text(encoding="utf-8"))
except (json.JSONDecodeError, OSError):
return {}
def _write_secrets(secrets: dict):
WEBHOOK_SECRETS_FILE.parent.mkdir(parents=True, exist_ok=True)
tmp = WEBHOOK_SECRETS_FILE.with_suffix(".tmp")
tmp.write_text(json.dumps(secrets, indent=2), encoding="utf-8")
tmp.replace(WEBHOOK_SECRETS_FILE)
try:
WEBHOOK_SECRETS_FILE.chmod(0o600)
except OSError:
pass # Windows doesn't support Unix permissions
def _store_secret(wh_id: str, secret: str | None) -> None:
secrets = _read_secrets()
if secret:
secrets[wh_id] = secret
else:
secrets.pop(wh_id, None)
_write_secrets(secrets)
def _get_secret(wh: dict) -> str | None:
"""Resolve a webhook secret from env, dedicated store, or legacy record."""
env_key = "OBSIGATE_WEBHOOK_SECRET_" + wh["id"].replace("-", "_").upper()
env_val = os.environ.get(env_key)
if env_val:
return env_val
stored = _read_secrets().get(wh["id"])
if stored:
return stored
return wh.get("secret") # legacy inline secret
def _public_view(wh: dict) -> dict:
"""Return a webhook record safe to expose through the API."""
clean = {k: v for k, v in wh.items() if k != "secret"}
clean["has_secret"] = bool(_get_secret(wh))
return clean
def get_webhooks() -> list:
return _read()
return [_public_view(wh) for wh in _read()]
def create_webhook(name: str, url: str, events: list[str], secret: str | None = None) -> dict:
validate_webhook_url(url)
webhooks = _read()
wh_id = str(uuid.uuid4())
wh = {
"id": str(uuid.uuid4()),
"id": wh_id,
"name": name,
"url": url,
"events": [e for e in events if e in VALID_EVENTS],
"secret": secret,
"enabled": True,
"created_at": datetime.now(timezone.utc).isoformat(),
"last_fired_at": None,
}
webhooks.append(wh)
_write(webhooks)
if secret:
_store_secret(wh_id, secret)
logger.info(f"Created webhook '{name}' → {url}")
return wh
return _public_view(wh)
def update_webhook(wh_id: str, updates: dict) -> dict | None:
webhooks = _read()
for wh in webhooks:
if wh["id"] == wh_id:
wh.update({k: v for k, v in updates.items() if k != "id"})
if updates.get("url"):
validate_webhook_url(updates["url"])
if "secret" in updates:
_store_secret(wh_id, updates["secret"])
safe_updates = {
k: v for k, v in updates.items()
if k not in ("id", "secret")
}
wh.update(safe_updates)
_write(webhooks)
return wh
return _public_view(wh)
return None
@@ -83,6 +232,9 @@ def delete_webhook(wh_id: str) -> bool:
if len(new_list) == len(webhooks):
return False
_write(new_list)
secrets = _read_secrets()
if secrets.pop(wh_id, None) is not None:
_write_secrets(secrets)
return True
@@ -102,17 +254,24 @@ async def dispatch_webhooks(event_type: str, data: dict):
async def _post(wh):
try:
# BUG-026: re-check the target right before connecting (DNS rebinding).
if not is_safe_target(wh["url"]):
logger.warning(f"Webhook '{wh['name']}' blocked by SSRF policy")
return
headers = {"Content-Type": "application/json", "X-ObsiGate-Event": event_type}
if wh.get("secret"):
sig = hmac.new(wh["secret"].encode(), body.encode(), hashlib.sha256).hexdigest()
secret = _get_secret(wh)
if secret:
sig = hmac.new(secret.encode(), body.encode(), hashlib.sha256).hexdigest()
headers["X-ObsiGate-Signature"] = f"sha256={sig}"
timeout = aiohttp.ClientTimeout(total=5)
async with aiohttp.ClientSession(timeout=timeout) as session, session.post(wh["url"], data=body, headers=headers) as resp:
if resp.status < 400:
logger.debug(f"Webhook '{wh['name']}' OK ({resp.status})")
else:
logger.warning(f"Webhook '{wh['name']}' failed ({resp.status})")
async with aiohttp.ClientSession(timeout=timeout) as session, session.post(
wh["url"], data=body, headers=headers, allow_redirects=False
) as resp:
if resp.status < 400:
logger.debug(f"Webhook '{wh['name']}' OK ({resp.status})")
else:
logger.warning(f"Webhook '{wh['name']}' failed ({resp.status})")
update_webhook(wh["id"], {"last_fired_at": datetime.now(timezone.utc).isoformat()})
except Exception as e:
logger.warning(f"Webhook '{wh['name']}' error: {e}")
+28 -1
View File
@@ -14,7 +14,7 @@
- **Projet** : ObsiGate — Porte d'entrée web pour vaults Obsidian
- **Stack** : Python 3.11+ (backend FastAPI) · JavaScript/Vanilla (frontend) · Tauri/Rust (desktop)
- **Dernière mise à jour** : 2026-09-12
- **Dernière mise à jour** : 2026-09-13
---
@@ -126,6 +126,30 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-014* | Assistant IA : les flèches ↑/↓ ne naviguent pas dans les menus `/` et `@` | 🟢 corrigé | P1 | 📱 frontend | IA | `frontend/js/bookslm.js` | Ouvrir `/`, puis ↑/↓ (y compris après un clic hors zone de texte) | Navigation gérée au niveau du **panneau en phase de capture** (`panel.addEventListener('keydown', ..., true)`), donc indépendante du focus | Le `keydown` n'était écouté que sur la zone de texte ; dès que le focus changeait, les flèches étaient ignorées. Tests : `tests/frontend/ai.test.mjs` (+1) |
| *BUG-015* | Assistant IA : la ligne sélectionnée des menus `/` et `@` est invisible au clavier | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/style.css` | Ouvrir `/` ou `@`, naviguer avec ↑/↓ | État `.active` en `--bg-hover` + barre d'accent à gauche (`inset 3px 0 0 var(--accent)`) au lieu de `--surface2` | `--surface2` est identique à `--bg-primary` (fond du menu) en thème sombre : la sélection ne se voyait pas. Idem pour la liste de modèles |
| *BUG-016* | Mobile : le bouton « mode lecture » flottant recouvre le bouton d'envoi de l'assistant IA | 🟢 corrigé | P2 | 📱 frontend | IA | `frontend/js/mobile-editor.js`, `frontend/style.css` | Mobile, sans fichier ouvert : ouvrir l'assistant puis constater que `📖` masque `✈️` | `updateReadingButtonVisibility()` n'affiche `#me-reading-btn` que si `state.currentPath` est défini **et** que le panneau assistant est fermé ; resynchronisé à chaque mutation de `#content-area` et sur `bookslm:opened`/`bookslm:closed` ; règle `.me-reading-btn[hidden]{display:none}` | Le bouton flottant (z-index 890) passait au-dessus du panneau plein écran mobile (z-index 100). Tests : `tests/frontend/mobile-editor.test.mjs` (+2) |
| *BUG-017* | Accueil mobile : les tuiles de « Favoris » et « Récents » ne s'affichent pas correctement en largeur (l'onglet « Partagés » s'affiche correctement) | 🔴 ouvert | P2 | 📱 frontend | IA | `frontend/js/dashboard.js`, `frontend/js/viewer.js`, `frontend/style.css` | Mobile : page d'accueil → onglets Favoris / Récents | — | Grille de tuiles non adaptée au viewport (largeur / nombre de colonnes) ; l'onglet « Partagés » (`.shared-card`) est conforme |
| *BUG-018* | Éditeur mobile : mise en page inadaptée — boutons « Annuler » / « Sauvegarder » inaccessibles et barre d'outils IA non optimisée | 🔴 ouvert | P1 | 📱 frontend | IA | `frontend/index.html`, `frontend/style.css`, `frontend/js/ai.js` | Mobile : ouvrir un fichier → barre d'outils de l'éditeur puis panneau assistant | — | Boutons de l'éditeur hors écran / masqués sur petit écran ; barre d'outils IA trop dense ou mal positionnée |
| *BUG-019* | Éditeur mobile : la barre d'outils du bas n'est pas toujours visible au-dessus du clavier quand le curseur n'est pas en bas de l'éditeur | 🔴 ouvert | P1 | 📱 frontend | IA | `frontend/js/mobile-editor.js`, `frontend/style.css` | Mobile : ouvrir un fichier, placer le curseur au milieu du document, afficher le clavier | — | La barre ne suit pas la position du curseur / n'est pas ancrée au-dessus du clavier virtuel |
| *BUG-020* | Éditeur mobile : deux barres de défilement verticales à droite de la page en mode édition | 🔴 ouvert | P3 | 📱 frontend | IA | `frontend/style.css` | Mobile : ouvrir l'éditeur et observer le bord droit | — | Double scrollbar (conteneur de l'éditeur + `.cm-scroller`) peu naturelle pour un éditeur |
| *BUG-021* | [🔴 CRITIQUE] XSS stocké via le rendu markdown (`escape=False`) | 🟢 corrigé | P0 | 🔐 sécurité | IA | `backend/services/sanitizer.py`, `backend/main.py` | Ouvrir une note contenant `<img src=x onerror=alert(1)>` | `backend/services/sanitizer.py` (whitelist stdlib) appliqué après `_add_heading_ids` dans `_render_markdown` ; tests `tests/test_security_hardening.py` | HTML brut d'un vault injecté dans le DOM → JS dans le navigateur de chaque utilisateur. DOMPurify client non ajouté (défense serveur suffisante) |
| *BUG-022* | [🔴 CRITIQUE] XSS stocké sur la page publique `/s/{token}` (title + frontmatter non échappés) | 🟢 corrigé | P0 | 🔐 sécurité | IA | `backend/main.py` | Partager une note dont le frontmatter contient un `title` avec `<script>` | `backend/main.py` : `html.escape()` sur `title`/frontmatter + JSON échappé (`\u003c`) pour le bloc `<script>` et `a.download` | `a.download="{title}.md"` se trouvait dans un bloc script → breakout JS. Test d'intégration dans `tests/test_security_hardening.py` |
| *BUG-023* | [🔴 CRITIQUE] Brute-force MFA sans rate-limit ni verrouillage | 🟢 corrigé | P0 | 🔐 sécurité | IA | `backend/auth/router.py` | `POST /api/auth/mfa/totp/verify` en boucle avec des codes aléatoires | `_enforce_mfa_rate_limit` + `_record_mfa_failure` sur `totp/verify`, `recovery`, `webauthn/verify` (IP + compte + lockout) | TOTP 6 chiffres brute-forceable ; le login était protégé, pas le second facteur |
| *BUG-024* | [🔴 CRITIQUE] Traversal inter-vaults : `startswith()` sans séparateur dans `resolve_safe_path` | 🟢 corrigé | P0 | 🔐 sécurité | IA | `backend/services/paths.py` | Avec les vaults `vault` et `vault-evil`, lire un fichier de `vault-evil` via `vault` | `_is_within()` par segments (`relative_to` + repli casse-insensible) ; tests `tests/test_security_hardening.py::TestPathIsolation` | Rupture d'isolation entre vaults (lecture / écriture / suppression) |
| *BUG-025* | [🟡 IMPORTANT] ReDoS : regex utilisateur appliquée au contenu en masse | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/services/regex_safety.py`, `backend/search.py`, `backend/services/mutations.py` | Recherche avancée avec `^(a+)+$` sur un gros document | Validation de pattern (longueur ≤500, rejet quantificateurs imbriqués/backrefs), contenu tronqué (200k) et matchs plafonnés ; 400 sur pattern refusé | Le dry-run `replace_in_files` déclenchait la charge ; la validation est faite avant compilation |
| *BUG-026* | [🟡 IMPORTANT] Webhooks : SSRF (URL non validée) + secret en clair | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/webhooks.py`, `data/webhook_secrets.json` | Configurer une URL vers `http://169.254.169.254` puis déclencher le webhook | Validation d'URL (HTTPS par défaut, IP privées/boucle bloquées), résolution DNS au dispatch, redirections non suivies ; secret dans `webhook_secrets.json` (0600) ou env `OBSIGATE_WEBHOOK_SECRET_<ID>` | `get_webhooks()` n'expose plus le secret (`has_secret`) ; options `OBSIGATE_WEBHOOK_ALLOW_HTTP`/`_ALLOW_PRIVATE` |
| *BUG-027* | [🟡 IMPORTANT] Sessions : refresh non rotatif, access token non révoqué au logout | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/auth/jwt_handler.py`, `backend/auth/router.py`, `backend/auth/middleware.py` | Logout puis réutilisation de l'ancien access token | Rotation du refresh token (`create_refresh_token(remember=…)`) + révocation de l'ancien JTI ; access token révoqué au logout et vérifié dans le middleware | JTI persistés dans `data/revoked_tokens.json` |
| *BUG-028* | [🟡 IMPORTANT] Politique de mot de passe incohérente + sessions non invalidées au changement | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/auth/password.py`, `backend/auth/router.py`, `backend/auth/user_store.py`, `backend/auth/middleware.py` | Créer un utilisateur avec `password:""` (admin) ; changer son mot de passe puis réutiliser l'ancien token | `validate_password_strength` (8–128) sur création/modif admin/changement ; `password_changed_at` invalide les jetons émis avant le changement ; `change-password` réémet une paire | Tests `TestPasswordPolicy` + `TestTokenInvalidation` |
| *BUG-029* | [🟡 IMPORTANT] Race read-modify-write sur `users.json` (perte de mises à jour) | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/auth/user_store.py` | Activer MFA et changer son mot de passe en parallèle | `threading.RLock` (`_users_lock`) autour de `create_user`, `update_user`, `delete_user`, `record_login_failure` | RLock réentrant car `record_login_failure` appelle `update_user` |
| *BUG-030* | [🟡 IMPORTANT] Adresse IP jamais consignée dans les audits (`_request_ip` toujours « unknown ») | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/services/net.py`, `backend/auth/middleware.py` | Lire le journal d'audit après un save/delete | `get_client_ip()` (`request.client.host`, `X-Forwarded-For` si `OBSIGATE_TRUST_PROXY=true`) injecté dans `current_user["_request_ip"]` | Test `TestClientIp` |
| *BUG-031* | [🟡 IMPORTANT] Rate-limiter en mémoire, mono-process, budget exclusivement IP | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/ratelimit.py`, `backend/auth/router.py` | Lancer plusieurs workers et observer le partage du compteur | Budget **par compte** (`is_account_rate_limited`/`record_account_*`) en plus de l'IP, intégré login + MFA ; limite mono-process documentée | Rotation d'IP neutralisée ; stockage partagé (Redis) hors périmètre |
| *BUG-032* | [🟡 IMPORTANT] Indexation : symlinks suivis (contenu hors vault indexé) + scan initial coûteux | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/indexer.py` | Placer un symlink dans le vault vers un dossier externe puis relancer l'index | `_scan_vault` réécrit avec `os.walk(followlinks=False)` + refus des symlinks sortant de la racine ; test `TestSymlinkIndexing` | Scan incrémental/index persistant : voir #86 (phase 3) |
| *BUG-033* | [🟡 IMPORTANT] Recherche classique et tool IA `search_fulltext` en O(N) sans inverted index | 🟢 corrigé | P1 | ⚙️ backend | IA | `backend/search.py`, `backend/tools/service.py` | `GET /api/search` sur un vault de 50 000 fichiers | `search()` récupère les candidats via l'inverted index (intersection des termes + expansion de préfixes), repli sur le scan pendant la construction | `search_fulltext` en bénéficie automatiquement |
| *BUG-034* | [🟡 IMPORTANT] CSP affaiblie (`'unsafe-inline'` + CDN distants) et token d'accès en sessionStorage | 🟢 corrigé | P1 | 🔐 sécurité | IA | `backend/main.py`, `frontend/js/auth.js`, `frontend/js/admin.js`, `frontend/js/sync.js` | Inspecter les en-têtes CSP ; lire sessionStorage en console | Token en mémoire + cookie HttpOnly (plus de `sessionStorage`) ; CSP durcie (`object-src 'none'`, `base-uri`, `form-action`, `frame-ancestors`). *Reste : migration nonce* | `'unsafe-inline'` conservé tant que les gestionnaires inline n'ont pas été convertis (résidu documenté) |
| *BUG-035* | [🔵 MINEUR] `secret_redactor` : faux positifs sur les hashs hex (git, SHA) | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/secret_redactor.py` | Lire une note contenant un commit git (40 caractères hexadécimaux) | Restreindre le périmètre de détection (contexte clé/token) + whitelist | Contenus mutilés dans les lectures et réponses IA |
| *BUG-036* | [🔵 MINEUR] Collab WebSocket : token en query string | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/collab.py` | Observer l'URL du websocket dans le trafic réseau | Passer le token en header / étape d'authentification initiale ; borner la taille des messages | Jeton visible dans les logs/proxys |
| *BUG-037* | [🔵 MINEUR] Compte « anonymous » administrateur si auth désactivée | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/middleware.py` | Démarrer avec l'authentification désactivée | Avertissement explicite au démarrage + refus de déploiement public sans auth | Comportement par conception mais risqué si mal configuré |
| *BUG-038* | [🔵 MINEUR] Argon2 à 64 MB par vérification : risque d'épuisement mémoire | 🔴 ouvert | P2 | 🔐 sécurité | IA | `backend/auth/password.py:8` | Lancer de nombreux `POST /api/auth/login` simultanés | Recalibrer (~19 MB, t=2, p=1, norme OWASP actuelle) + maintien du rate-limit | DoS mémoire possible sur les petites instances |
| *BUG-039* | [🔵 MINEUR] Enumération de comptes : 429 (verrouillé) vs 401 (inconnu) | 🔴 ouvert | P3 | 🔐 sécurité | IA | `backend/auth/router.py:120` | Tenter un login sur un compte verrouillé puis un nom inconnu | Répondre 401 uniforme avec un timing équivalent | Le statut HTTP distingue l'existence d'un compte |
| *BUG-040* | [🔵 MINEUR] Extraction PDF intégrale (100 ko) au scan de démarrage | 🔴 ouvert | P2 | ⚙️ backend | IA | `backend/indexer.py:465` | Démarrer sur un vault contenant de nombreux PDF | Analyser les PDF en tâche de fond / à la demande (lazy) | Ralentit fortement le démarrage et le rebuild d'index |
### TODOs techniques (améliorations / nouvelles tâches)
@@ -158,6 +182,9 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| 2026-09-12 | #82 | Amélioration perf + UX | `backend/services/search.py`, `backend/main.py`, `frontend/js/bookslm.js`, `frontend/js/ui.js`, `frontend/style.css`, `tests/test_api_main.py`, `tests/frontend/ai.test.mjs` | #82 : endpoint `GET /api/vault/{vault}/paths` + préchargement et filtrage client du menu `@` (instantané) ; sélecteurs Fournisseur/Modèle agrandis (0,8 rem / 34 px) ; « 🧠 BooksLM » ajouté au menu contextuel de la racine des vaults. Vérifié : pytest 151 (ciblés) + ruff/mypy, tests frontend 49/49. | ✅ livré (en attente vérif utilisateur) |
| 2026-09-12 | #82 | UX | `frontend/js/ai.js`, `frontend/style.css`, `tests/frontend/ai.test.mjs` | Sélecteurs Fournisseur/Modèle alignés à droite dans la barre de l'assistant et réordonnés : capacité du modèle (ⓘ) → fournisseur → modèle. Bulle de capacités ouverte vers la droite (`left: 0`). Test JSDOM de l'ordre. Vérifié : tests frontend 49/49. | ✅ livré (en attente vérif utilisateur) |
| 2026-09-12 | BUG-016 | Correction | `frontend/js/mobile-editor.js`, `frontend/style.css`, `tests/frontend/mobile-editor.test.mjs`, `tests/e2e/mobile-editor.spec.js` | BUG-016 : le bouton « mode lecture » n'apparaît que si un fichier est ouvert et que l'assistant est fermé, sinon il recouvrait le bouton d'envoi sur mobile. Vérifié : tests frontend 24/24 (mobile) + suite JSDOM verte. | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-13 | BUG-017, BUG-018, BUG-019, BUG-020, #83 | Enregistrement | `docs/ISSUES_TODOLIST.md`, `docs/ROADMAP.md` | Signalements mobile consignés : accueil (tuiles Favoris/Récents), éditeur (boutons Annuler/Sauvegarder inaccessibles + barre IA), visibilité de la barre d'outils au-dessus du clavier, double barre de défilement. Feature #83 ajoutée : barre d'outils d'édition mobile style Obsidian Android. | 🔴 ouvert (à traiter) |
| 2026-09-13 | Revue sécurité statique (BUG-021 → BUG-040, roadmap #84 → #87) | Enregistrement | `docs/ISSUES_TODOLIST.md`, `docs/ROADMAP.md` | Revue statique 2026-09-13 consignée : XSS markdown (`escape=False`) + page de partage publique, brute-force MFA, traversal par préfixe `paths.py`, ReDoS, SSRF webhooks, cycle de vie des sessions, politique de mot de passe, races `users.json`, audits IP, rate-limit, indexation symlinks, recherche O(N), CSP. BUG-021→BUG-040 ouverts au registre ; roadmap #84→#87 (consolidation/sécurité, refonte architecturale, performance, CI/CD). | 🔴 ouvert (à traiter) |
| 2026-09-13 | BUG-021 → BUG-034 (#84 phase 1) | Correction | `backend/services/sanitizer.py`, `backend/services/paths.py`, `backend/services/net.py`, `backend/services/regex_safety.py`, `backend/main.py`, `backend/search.py`, `backend/webhooks.py`, `backend/indexer.py`, `backend/ratelimit.py`, `backend/auth/{router,user_store,password,jwt_handler,middleware}.py`, `frontend/js/{auth,admin,sync}.js`, `tests/test_security_hardening.py`, `tests/test_auth_api.py` | Sanitizer XSS serveur (markdown + page de partage), rate-limit/lockout MFA, isolation vaults par segments, caps regex (ReDoS), SSRF webhooks + secrets externalisés, rotation/révocation des jetons, politique de mot de passe + invalidation des sessions, verrous `users.json`, IP réelle dans les audits, rate-limit par compte, symlinks d'index ignorés, recherche simple via inverted index, token en mémoire + cookie HttpOnly + CSP durcie. Vérifié : pytest 961 passed / 6 skipped, ruff 0, mypy 0, tests frontend 36 modules + 9 JSDOM suites verts. | 🟢 corrigé (en attente vérif utilisateur) |
---
+71 -2
View File
@@ -1,6 +1,6 @@
# ObsiGate — Roadmap
> **Version :** 2.3.0-dev | **Dernière mise à jour :** 2026-09-12
> **Version :** 2.3.0-dev | **Dernière mise à jour :** 2026-09-13
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
> vers les fonctionnalités livrées.
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
@@ -60,6 +60,72 @@
---
## ⚪ Backlog — Priorité 2 (P2)
### 83. Barre d'outils d'édition mobile — style Obsidian Android
- **Effort :** 3-5 jours | **Impact :** 🟡 | **Zone :** frontend (mobile)
- **Description :** remplacer la barre de mise en forme Markdown actuelle par un **ruban horizontal
défilable** ancré juste au-dessus du clavier virtuel, reprenant l'ergonomie de l'app Android
Obsidian : fond anthracite aux coins arrondis, insertion/enrobage de la syntaxe au curseur ou sur
la sélection, et personnalisation des commandes via une icône clé à molette.
- **Sous-tâches :**
- [ ] Ruban horizontal défilable (glissement tactile gauche/droite) ancré au-dessus du clavier
- [ ] Actions rapides : annuler, refaire, `[[ ]]` (lien interne), modèle/fichiers, tag `#`, pièce jointe
- [ ] Formatage : H1–H6, gras, italique, barré (`~~`), surligné (`==`), code en ligne/bloc, citation (`>`)
- [ ] Liens externes, listes à puces/numérotées, case à cocher (`- [ ]`), indenter / désindenter
- [ ] Personnalisation (clé à molette) : ajouter / supprimer / réordonner les commandes
- [ ] i18n FR/EN + tests frontend (helpers purs) + E2E mobile
---
## ⚪ Backlog — Sécurité, architecture & performance (P0/P1)
### 84. Consolidation & sécurité — revue statique 2026-09-13 (phase 1)
- **Effort :** 6-9 jours | **Impact :** 🔴 | **Zone :** backend + frontend | **Référence :** [ISSUES_TODOLIST.md](./ISSUES_TODOLIST.md) BUG-021 → BUG-034
- **Statut :** 🟢 livré (phase 1) — sanitizer XSS, rate-limit/lockout MFA, isolation vaults, ReDoS, SSRF webhooks, cycle de vie des sessions, politique de mot de passe, verrous `users.json`, audits IP, rate-limit par compte, symlinks, recherche via inverted index, token en cookie HttpOnly. Détail : [archive/COMPLETED_v1-v2.md](./archive/COMPLETED_v1-v2.md) (section #84).
- **Description :** traiter toutes les vulnérabilités critiques et importantes issues de la revue statique : XSS markdown (`escape=False`) et page publique de partage, brute-force MFA, isolation des vaults (`resolve_safe_path`), ReDoS, SSRF webhooks, cycle de vie des sessions, politique de mot de passe, races `users.json`, audits IP, rate-limit partagé, indexation symlinks.
- **Sous-tâches :**
- [x] Assainir le rendu markdown (sanitizer serveur en whitelist) et la page de partage (échappement `title`/frontmatter) — *DOMPurify client non ajouté (défense en profondeur serveur suffisante)*
- [x] Rate-limit + lockout sur les endpoints MFA (`totp/verify`, `recovery`, `webauthn/verify`)
- [x] Corriger `resolve_safe_path` (comparaison de chemin stricte par segment) + test de régression
- [x] Rotation du refresh token, révocation de l'access token au logout, persistance des JTI révoqués
- [x] Valider la politique de mot de passe à la création ; bloquer le SSRF des webhooks et externaliser les secrets
- [x] Verrous sur les mutations `users.json` ; consigner l'adresse IP réelle dans les audits
- [x] Ignorer les symlinks de l'index ; caps CPU/timeout regex (ReDoS)
- [~] Durcir la CSP — *partiel* : directives `object-src`/`base-uri`/`form-action`/`frame-ancestors` ajoutées et token retiré de `sessionStorage` ; migration **nonce** restante (nécessite la conversion des gestionnaires d'événements inline)
### 85. Refonte architecturale — découpage du monolithe & persistance d'état (phase 2)
- **Effort :** 8-12 jours | **Impact :** 🟡 | **Zone :** backend
- **Description :** extraire le monolithe `backend/main.py` (~4 260 lignes) en routers FastAPI par domaine et rendre persistant l'état qui ne l'est pas (index de recherche, JTI révoqués, compteurs de rate-limit) pour préparer le multi-nœuds.
- **Sous-tâches :**
- [ ] Routers par domaine : files, search, share, webhooks, plugins, collab, admin, ai
- [ ] Centraliser le contrat d'outils IA sur `tools/registry.py` (permissions, quotas, redaction)
- [ ] Persister index, JTI révoqués et compteurs de rate-limit (SQLite/Redis)
- [ ] Verrous asyncio autour de l'index global et des stores JSON ; service de partage public (expiration, révocation, quotas)
### 86. Optimisation globale des performances (phase 3)
- **Effort :** 4-6 jours | **Impact :** 🟡 | **Zone :** backend (`search.py`, `indexer.py`, `mutations.py`)
- **Description :** brancher l'inverted index (déjà construit) sur la recherche simple et le tool IA `search_fulltext`, indexation incrémentale, extraction PDF lazy.
- **Sous-tâches :**
- [ ] Recherche simple + tool IA via l'inverted index (suppression du balayage O(N) en mémoire)
- [ ] Indexation incrémentale + scan différentiel au démarrage (remplace le `rglob` complet)
- [ ] Extraction PDF/excalidraw différée (hors scan) ; caps CPU sur les opérations regex
### 87. Amélioration continue — tests, CI/CD, revues de sécurité (phase 4)
- **Effort :** 3-5 jours | **Impact :** 🟡 | **Zone :** `.gitea/workflows/`, `tests/`
- **Description :** renforcer le pipeline (`.gitea/workflows/ci.yml`, `desktop-build.yml`) pour le rendre bloquant par défaut et accompagner les phases 1 → 3.
- **Sous-tâches :**
- [ ] Jobs CI sécurité (bandit/semgrep/trivy, audits pip/npm) + tests E2E XSS (page de partage + lecteur markdown)
- [ ] Tests de concurrence (`users.json`), fuzzing de timing regex, couverture des composants critiques
- [ ] Revue périodique des dépendances ; documentation utilisateur FR/EN synchronisée ; contrôle automatisé de la conformité au DoD
---
## ✅ Complété — index
> Détail complet dans [docs/archive/COMPLETED_v1-v2.md](./archive/COMPLETED_v1-v2.md) et
@@ -99,6 +165,7 @@
| 69 | Éditeur mobile natif — Interface tactile optimisée | 2.3.0 | [features/mobile-editor.md](./features/mobile-editor.md) |
| 70 | Recherche sémantique — Embeddings vectoriels (hybride TF-IDF + RRF) | 2.3.0 | [features/semantic-search.md](./features/semantic-search.md) |
| 77 | Application Desktop native — Tauri | 🔵 en cours | [features/desktop-tauri.md](./features/desktop-tauri.md) |
| 84 | Consolidation & sécurité — revue statique 2026-09-13 (phase 1, BUG-021→034) | 2.3.0 | [archive](./archive/COMPLETED_v1-v2.md) |
---
@@ -108,8 +175,10 @@
|---|---|---|
| ✅ Complété | #1 → #59, #61–72, #74–76, #78–82 | ~107 jours réalisés |
| 🔵 P2 restant | #77 Desktop : signature de code (non retenue), 6 tests E2E **manuels** ([protocole](./DESKTOP_E2E_CHECKLIST.md)) | ~0,5-1 jour |
| ⚪ P2 restant | #83 Barre d'outils d'édition mobile style Obsidian (3-5j) | 3-5 jours |
| ⚪ P4 restant | #73 Sync (6-8j) | 6-8 jours |
| **Total restant** | **2 items + finitions** | **~7-10 jours** |
| ⚪ P0/P1 restant | #85-87 Refonte architecturale, performance, CI/CD (issues BUG-035 → BUG-040) | ~15-23 jours |
| **Total restant** | **7 items + finitions** | **~31-47 jours** |
---
+25
View File
@@ -317,6 +317,31 @@
---
## #84 — Consolidation & sécurité (phase 1) ✅ TERMINÉ
Traitement des vulnérabilités de la revue statique du 2026-09-13 (BUG-021 → BUG-034).
- **Sanitizer XSS serveur** (`backend/services/sanitizer.py`, stdlib, liste blanche
de balises/attributs/schémas) appliqué au rendu markdown et à la page publique
`/s/{token}` (titre, frontmatter et JSON échappés, `</script>` neutralisé).
- **Auth durcie** : rate-limit IP + compte et verrouillage sur les endpoints MFA ;
rotation du refresh token et révocation de l'access token au logout ; politique de
mot de passe centralisée (8–128) ; `password_changed_at` invalide les jetons
antérieurs ; verrou `RLock` sur `users.json`.
- **Isolation & réseau** : `resolve_safe_path` compare par segments (`vault` ≠
`vault-evil`) ; webhooks protégés contre le SSRF (HTTPS, IP privées bloquées,
résolution vérifiée, redirections interdites, secret externalisé) ; IP client
réelle dans les audits.
- **Robustesse** : validation des regex (ReDoS), symlinks hors vault ignorés à
l'indexation, recherche simple et `search_fulltext` branchés sur l'inverted index.
- **Frontend** : token d'accès en mémoire + cookie `HttpOnly` (plus de
`sessionStorage`), CSP durcie (`object-src`, `base-uri`, `form-action`,
`frame-ancestors`).
- **Tests** : `tests/test_security_hardening.py` (30) ; suite backend 961 passed /
6 skipped ; ruff + mypy 0 ; suites frontend vertes.
---
## Grosses fonctionnalités — fiches dédiées
| # | Feature | Version | Fiche |
+4 -9
View File
@@ -28,17 +28,12 @@ let _auditFilters = { user: "", action: "" };
// ── Helpers ──────────────────────────────────────────────────────────────
/**
* Read the current Bearer token from sessionStorage (same key auth.js uses).
* Returns an `{Authorization: "Bearer ..."}` object or null.
* BUG-034: the access token now lives in an HttpOnly cookie (and in memory
* inside auth.js), never in sessionStorage. Same-origin requests carry the
* cookie automatically, so no Authorization header is needed here.
*/
export function getAuthHeaders() {
try {
const token = sessionStorage.getItem("obsigate_access_token");
if (!token) return null;
return { Authorization: "Bearer " + token };
} catch {
return null;
}
return null;
}
/** Format a size in MB as a human-readable string. */
+18 -6
View File
@@ -85,12 +85,18 @@ const AuthManager = {
TOKEN_EXPIRY_KEY: "obsigate_token_expiry",
USER_KEY: "obsigate_user",
_authEnabled: false,
// BUG-034: the access token is kept in memory only. The server also sets it
// as an HttpOnly cookie, so a page reload re-authenticates via /api/auth/refresh
// without ever exposing the token to JavaScript-readable storage.
_accessToken: null,
// ── Token storage (sessionStorage) ─────────────────────────────
// ── Token storage (in-memory + HttpOnly cookie) ────────────────
saveToken(tokenData) {
const expiresAt = Date.now() + tokenData.expires_in * 1000;
sessionStorage.setItem(this.ACCESS_TOKEN_KEY, tokenData.access_token);
this._accessToken = tokenData.access_token;
// Clear any token persisted by an older build (XSS-readable).
try { sessionStorage.removeItem(this.ACCESS_TOKEN_KEY); } catch (e) { /* ignore */ }
sessionStorage.setItem(this.TOKEN_EXPIRY_KEY, expiresAt.toString());
if (tokenData.user) {
sessionStorage.setItem(this.USER_KEY, JSON.stringify(tokenData.user));
@@ -98,7 +104,11 @@ const AuthManager = {
},
getToken() {
return sessionStorage.getItem(this.ACCESS_TOKEN_KEY);
return this._accessToken;
},
hasSession() {
return !!this._accessToken || !!sessionStorage.getItem(this.TOKEN_EXPIRY_KEY);
},
getUser() {
@@ -114,7 +124,8 @@ const AuthManager = {
},
clearSession() {
sessionStorage.removeItem(this.ACCESS_TOKEN_KEY);
this._accessToken = null;
try { sessionStorage.removeItem(this.ACCESS_TOKEN_KEY); } catch (e) { /* ignore */ }
sessionStorage.removeItem(this.TOKEN_EXPIRY_KEY);
sessionStorage.removeItem(this.USER_KEY);
},
@@ -281,7 +292,8 @@ const AuthManager = {
}
const data = await response.json();
const expiry = Date.now() + data.expires_in * 1000;
sessionStorage.setItem(this.ACCESS_TOKEN_KEY, data.access_token);
this._accessToken = data.access_token;
try { sessionStorage.removeItem(this.ACCESS_TOKEN_KEY); } catch (e) { /* ignore */ }
sessionStorage.setItem(this.TOKEN_EXPIRY_KEY, expiry.toString());
return data.access_token;
},
@@ -359,7 +371,7 @@ const AuthManager = {
}
// Auth enabled — check for existing session
if (this.getToken() && !this.isTokenExpired()) {
if (this.hasSession() && !this.isTokenExpired()) {
this.showApp();
return true;
}
+1 -1
View File
@@ -612,7 +612,7 @@ export function init() {
// fresh login, so we only do it here for already-authenticated or
// auth-disabled sessions.
setTimeout(() => {
if (!AuthManager._authEnabled || AuthManager.getToken()) {
if (!AuthManager._authEnabled || AuthManager.hasSession()) {
syncFileIndexFromServer();
}
}, 3000);
+17
View File
@@ -1168,6 +1168,23 @@ async function main() {
assert.ok(vaultBranch.includes("BooksLM"), "BooksLM entry present in the vault root context menu");
});
// ── BUG-034: the access token is not persisted in sessionStorage ──
await test("auth token kept in memory, not in sessionStorage", async () => {
const authMod = await import(pathToFileURL(path.join(JS_DIR, "auth.js")).href);
const { AuthManager } = authMod;
sessionStorage.clear();
AuthManager.saveToken({
access_token: "secret-token",
expires_in: 3600,
user: { username: "u", role: "user" },
});
assert.equal(sessionStorage.getItem("obsigate_access_token"), null, "token must not be stored");
assert.equal(AuthManager.getToken(), "secret-token", "token available in memory");
assert.ok(AuthManager.hasSession(), "session detected");
AuthManager.clearSession();
assert.equal(AuthManager.getToken(), null, "session cleared");
});
// ── Summary ──
console.log(`\n${passCount}/${testCount} tests passed`);
if (passCount !== testCount) {
+1 -1
View File
@@ -252,7 +252,7 @@ class TestAdmin:
"Authorization": f"Bearer {token}",
}, json={
"username": "deleteuser",
"password": "delpass",
"password": "delpass123",
"role": "user",
"vaults": ["TestVault"],
})
+1
View File
@@ -277,6 +277,7 @@ class TestWebhooks:
from backend import webhooks
self.wh_file = tmp_path / "webhooks.json"
monkeypatch.setattr(webhooks, "WEBHOOKS_FILE", self.wh_file)
monkeypatch.setattr(webhooks, "WEBHOOK_SECRETS_FILE", tmp_path / "webhook_secrets.json")
yield
def test_get_empty(self):
+319
View File
@@ -0,0 +1,319 @@
# tests/test_security_hardening.py — Regression tests for ROADMAP #84
# (BUG-021 → BUG-034): sanitizer, path isolation, password policy, regex
# safety, webhook SSRF, rate limiting, audit IP and share-page escaping.
import time
from pathlib import Path
import pytest
from backend.services.errors import ServiceError
# ═══════════════════════════════════════════════════════════════════
# BUG-021/022 — HTML sanitizer
# ═══════════════════════════════════════════════════════════════════
class TestSanitizer:
def test_strips_event_handlers(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<img src="x" onerror="alert(1)">')
assert "onerror" not in out
assert "alert" not in out
def test_drops_script_content(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html("<p>ok</p><script>alert(1)</script>tail")
assert "<script" not in out
assert "alert" not in out
assert "ok" in out
assert "tail" in out
def test_blocks_javascript_urls(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<a href="javascript:alert(1)">x</a>')
assert "javascript:" not in out
assert ">x</a>" in out
def test_blocks_obfuscated_javascript_url(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<a href="java&#115;cript:alert(1)">x</a>')
assert "javascript" not in out.lower()
def test_keeps_safe_markup(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<p class="a">hi</p><a href="/x" data-vault="v" data-path="p.md">y</a>')
assert '<p class="a">hi</p>' in out
assert 'data-vault="v"' in out
assert 'href="/x"' in out
def test_allows_image_data_uri_but_not_html(self):
from backend.services.sanitizer import sanitize_html
good = sanitize_html('<img src="data:image/png;base64,AAAA">')
assert "data:image/png" in good
bad = sanitize_html('<img src="data:text/html,<script>alert(1)</script>">')
assert "data:text/html" not in bad
def test_heading_ids_preserved(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<h2 id="intro">Intro</h2>')
assert 'id="intro"' in out
def test_drops_style_attribute(self):
from backend.services.sanitizer import sanitize_html
out = sanitize_html('<p style="position:fixed">x</p>')
assert "style" not in out
def test_is_safe_url(self):
from backend.services.sanitizer import is_safe_url
assert is_safe_url("/relative")
assert is_safe_url("https://example.com")
assert is_safe_url("#anchor")
assert not is_safe_url("javascript:alert(1)")
assert not is_safe_url("vbscript:msgbox")
assert not is_safe_url("data:text/html,<script>")
# ═══════════════════════════════════════════════════════════════════
# BUG-024 — Vault path isolation
# ═══════════════════════════════════════════════════════════════════
class TestPathIsolation:
def test_inside_vault_ok(self, tmp_path):
from backend.services.paths import resolve_safe_path
root = tmp_path / "vault"
root.mkdir()
resolved = resolve_safe_path(root, "note.md")
assert resolved == (root / "note.md").resolve()
def test_sibling_prefix_is_rejected(self, tmp_path):
from backend.services.paths import resolve_safe_path
(tmp_path / "vault").mkdir()
(tmp_path / "vault-evil").mkdir()
with pytest.raises(ServiceError) as exc:
resolve_safe_path(tmp_path / "vault", "../vault-evil/secret.md")
assert exc.value.code == "path_outside_vault"
def test_parent_traversal_rejected(self, tmp_path):
from backend.services.paths import resolve_safe_path
root = tmp_path / "vault"
root.mkdir()
with pytest.raises(ServiceError) as exc:
resolve_safe_path(root, "../../etc/passwd")
assert exc.value.code == "path_outside_vault"
# ═══════════════════════════════════════════════════════════════════
# BUG-028 — Password policy
# ═══════════════════════════════════════════════════════════════════
class TestPasswordPolicy:
def test_short_rejected(self):
from backend.auth.password import validate_password_strength
with pytest.raises(ValueError):
validate_password_strength("short")
def test_too_long_rejected(self):
from backend.auth.password import validate_password_strength
with pytest.raises(ValueError):
validate_password_strength("a" * 200)
def test_valid_accepted(self):
from backend.auth.password import validate_password_strength
assert validate_password_strength("ValidPass123") == "ValidPass123"
# ═══════════════════════════════════════════════════════════════════
# BUG-025 — Regex safety (ReDoS)
# ═══════════════════════════════════════════════════════════════════
class TestRegexSafety:
def test_nested_quantifier_rejected(self):
from backend.services.regex_safety import validate_regex
with pytest.raises(ValueError):
validate_regex("(a+)+$")
def test_long_pattern_rejected(self):
from backend.services.regex_safety import validate_regex
with pytest.raises(ValueError):
validate_regex("a" * 600)
def test_simple_pattern_accepted(self):
from backend.services.regex_safety import validate_regex
assert validate_regex(r"\bpython\b") == r"\bpython\b"
def test_truncate_for_regex(self):
from backend.services.regex_safety import truncate_for_regex
assert len(truncate_for_regex("x" * 1000, limit=10)) == 10
# ═══════════════════════════════════════════════════════════════════
# BUG-026 — Webhook SSRF
# ═══════════════════════════════════════════════════════════════════
class TestWebhookUrlValidation:
def test_https_public_ok(self):
from backend.webhooks import validate_webhook_url
assert validate_webhook_url("https://example.com/hook") == "https://example.com/hook"
def test_http_rejected_by_default(self, monkeypatch):
from backend.webhooks import validate_webhook_url
monkeypatch.delenv("OBSIGATE_WEBHOOK_ALLOW_HTTP", raising=False)
with pytest.raises(ValueError):
validate_webhook_url("http://example.com/hook")
def test_private_ip_rejected(self, monkeypatch):
from backend.webhooks import validate_webhook_url
monkeypatch.delenv("OBSIGATE_WEBHOOK_ALLOW_PRIVATE", raising=False)
with pytest.raises(ValueError):
validate_webhook_url("https://127.0.0.1/hook")
with pytest.raises(ValueError):
validate_webhook_url("https://169.254.169.254/latest/meta-data")
def test_bad_scheme_rejected(self):
from backend.webhooks import validate_webhook_url
with pytest.raises(ValueError):
validate_webhook_url("ftp://example.com")
# ═══════════════════════════════════════════════════════════════════
# BUG-031 — Per-account rate limiting
# ═══════════════════════════════════════════════════════════════════
class TestAccountRateLimit:
def test_account_limited_after_threshold(self, monkeypatch):
from backend import ratelimit
monkeypatch.setattr(ratelimit, "ACCOUNT_MAX_ATTEMPTS", 3)
account = "ratelimit-account@test"
assert not ratelimit.is_account_rate_limited(account)
for _ in range(3):
ratelimit.record_account_failure(account)
assert ratelimit.is_account_rate_limited(account)
ratelimit.record_account_success(account)
assert not ratelimit.is_account_rate_limited(account)
# ═══════════════════════════════════════════════════════════════════
# BUG-030 — Client IP resolution
# ═══════════════════════════════════════════════════════════════════
class TestClientIp:
def _request(self, headers=None, client=("1.2.3.4", 1234)):
from starlette.requests import Request
raw_headers = [(k.lower().encode(), v.encode()) for k, v in (headers or {}).items()]
scope = {
"type": "http",
"method": "GET",
"path": "/",
"query_string": b"",
"headers": raw_headers,
"client": client,
"scheme": "http",
"server": ("test", 80),
}
return Request(scope)
def test_socket_peer_used_by_default(self, monkeypatch):
from backend.services.net import get_client_ip
monkeypatch.delenv("OBSIGATE_TRUST_PROXY", raising=False)
assert get_client_ip(self._request()) == "1.2.3.4"
def test_forwarded_header_used_when_trusted(self, monkeypatch):
from backend.services.net import get_client_ip
monkeypatch.setenv("OBSIGATE_TRUST_PROXY", "true")
req = self._request({"X-Forwarded-For": "9.8.7.6, 10.0.0.1"})
assert get_client_ip(req) == "9.8.7.6"
# ═══════════════════════════════════════════════════════════════════
# BUG-032 — Indexing must not follow symlinks outside the vault
# ═══════════════════════════════════════════════════════════════════
class TestSymlinkIndexing:
def test_external_symlink_is_skipped(self, tmp_path):
import os
from backend.indexer import _scan_vault
vault = tmp_path / "vault"
vault.mkdir()
(vault / "inside.md").write_text("# Inside\n", encoding="utf-8")
outside = tmp_path / "outside"
outside.mkdir()
(outside / "secret.md").write_text("# Secret\n", encoding="utf-8")
link = vault / "link.md"
try:
os.symlink(outside / "secret.md", link)
except (OSError, NotImplementedError):
pytest.skip("Symlinks are not supported in this environment")
result = _scan_vault("vault", str(vault))
paths = {f["path"] for f in result["files"]}
assert "inside.md" in paths
assert "link.md" not in paths
# ═══════════════════════════════════════════════════════════════════
# BUG-021/022 — End-to-end rendering & share page escaping
# ═══════════════════════════════════════════════════════════════════
class TestRenderingAndShare:
def test_markdown_rendering_is_sanitized(self, client, test_vault_dir):
payload = (
"---\ntitle: Safe\n---\n"
"# Hello\n\n"
"<img src=x onerror=alert(1)>\n\n"
"<script>alert(2)</script>\n\n"
"Normal **text**.\n"
)
(Path(test_vault_dir) / "xss.md").write_text(payload, encoding="utf-8")
resp = client.get("/api/file/TestVault", params={"path": "xss.md"})
assert resp.status_code == 200
html = resp.json()["html"]
assert "onerror" not in html
assert "<script" not in html
assert "<strong>text</strong>" in html
def test_share_page_escapes_title_and_frontmatter(self, client, test_vault_dir):
payload = (
'---\ntitle: "<script>alert(1)</script>"\n'
'author: "<img src=x onerror=alert(9)>"\n---\n'
"# Body\ncontent\n"
)
(Path(test_vault_dir) / "share_xss.md").write_text(payload, encoding="utf-8")
create = client.post("/api/share/TestVault", json={"path": "share_xss.md"})
assert create.status_code == 200
token = create.json()["token"]
page = client.get(f"/s/{token}")
assert page.status_code == 200
body = page.text
assert "<script>alert(1)</script>" not in body
assert "&lt;script&gt;alert(1)&lt;/script&gt;" in body
# The img tag must be escaped, so no live attribute is emitted.
assert "<img src=x onerror" not in body
assert "&lt;img" in body
# ═══════════════════════════════════════════════════════════════════
# BUG-027/028 — Token invalidation on password change
# ═══════════════════════════════════════════════════════════════════
class TestTokenInvalidation:
def test_access_token_rejected_after_password_change(self, admin_client):
login = admin_client.post(
"/api/auth/login", json={"username": "admin", "password": "chab30"}
)
assert login.status_code == 200
token = login.json()["access_token"]
# Simulate a password change happening in the future relative to the token.
from backend.auth.user_store import update_user
update_user("admin", {"password_changed_at": time.time() + 10})
me = admin_client.get(
"/api/auth/me", headers={"Authorization": f"Bearer {token}"}
)
assert me.status_code == 401