Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6ba04c4381 | ||
|
|
06f8e63d06 |
+70
-1
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
|
||||
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
> **En cours de développement** : les changements à venir sont listés dans la section
|
||||
> [Unreleased](#unreleased). La dernière version livrée est **2.29.0**.
|
||||
> [Unreleased](#unreleased). La dernière version livrée est **2.31.0**.
|
||||
|
||||
---
|
||||
|
||||
@@ -14,6 +14,75 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
---
|
||||
|
||||
## [2.31.0] — 2026-09-28
|
||||
|
||||
### Correction
|
||||
|
||||
- **BUG-090 — troncature silencieuse d'une feuille `.xlsx` au-delà de
|
||||
500 lignes × 40 colonnes.** `render_sheets()` renvoie les dimensions
|
||||
déclarées par la feuille (`total_rows`/`total_cols`), les plafonds du
|
||||
moteur (`max_rows`/`max_cols`) et un flag `truncated` : la visionneuse
|
||||
affiche un bandeau « Feuille tronquée — 500 lignes affichées sur 520 »
|
||||
(i18n FR/EN) au lieu de présenter une table courte comme complète. La
|
||||
ligne d'en-têtes est désormais figée au défilement vertical (`thead`
|
||||
sticky, `top: auto` sur les numéros de ligne pour éviter leur
|
||||
empilement en haut à gauche). *#153 A8/R5.*
|
||||
|
||||
### Ajouté
|
||||
|
||||
- **#153 A9 — chargement paresseux d'une feuille par fenêtres.**
|
||||
`GET /api/file/{vault}/xlsx/sheet?sheet=&offset=&limit=` renvoie un
|
||||
bloc de lignes (`XlsxSheetWindowResponse`, plafond 1 000 lignes par
|
||||
requête, `has_more` de pagination) avec les **vraies** coordonnées A1
|
||||
et numéros de ligne de la feuille — une fenêtre se comporte exactement
|
||||
comme le rendu complet. Erreurs typées : 404 feuille inconnue, 415
|
||||
fichier non-`.xlsx`. La lecture des valeurs calculées en cache (#153
|
||||
A12) s'applique aussi aux fenêtres. Le défilement virtuel côté UI
|
||||
reste à faire ; l'endpoint rend les lignes au-delà du plafond déjà
|
||||
accessibles aux clients API.
|
||||
|
||||
---
|
||||
|
||||
## [2.30.0] — 2026-09-27
|
||||
|
||||
### Correction
|
||||
|
||||
- **BUG-089 — un reindex manuel ne reconstruisait pas l'index inversé.**
|
||||
`reload_index()` / `reload_single_vault()` remplacent l'entrée de vault
|
||||
en bloc, ce qui n'émet pas les notifications incrémentales : la
|
||||
recherche TF-IDF continuait de servir un index périmé après un
|
||||
reindex. Les deux fonctions appellent désormais `init_inverted_index()`.
|
||||
Au passage, `backend/search.py` lisait l'index via
|
||||
`from backend.indexer import index` — une liaison **par valeur** du
|
||||
dict : un rechargement du module `backend.indexer` recréait le dict
|
||||
côté indexer alors que la recherche écrivait dans l'ancien, et
|
||||
l'index inversé n'indexait plus rien. Tous les accès passent par
|
||||
`_indexer.index`. *Trouvé en écrivant le test de recherche d'A5 : il
|
||||
passait isolément et échouait en suite complète selon l'ordre.*
|
||||
|
||||
### Ajouté
|
||||
|
||||
- **#153 A5 — les tableurs sont indexés par leur contenu.**
|
||||
`extract_indexable_text()` extrait les noms de feuilles et les 20
|
||||
premières lignes (plafond 5 000 caractères, 20 feuilles) pour le
|
||||
TF-IDF et la recherche sémantique. Un mot tapé dans une cellule rend
|
||||
désormais le classeur trouvable ; la lecture binaire pour l'affichage
|
||||
est inchangée et un classeur chiffré/corrompu s'indexe par son seul
|
||||
nom.
|
||||
- **#153 A10 — la saisie est typée comme dans Excel.** Une valeur
|
||||
`TRUE`/`FAUX`/`OUI`/`NON` devient un booléen, une date `JJ/MM/AAAA`
|
||||
(avec `HH:MM` optionnel) devient une vraie date — et dans l'ordre
|
||||
français : `01/02/2026` est le 1ᵉʳ février. Une saisie ressemblant à
|
||||
une formule n'est jamais convertie.
|
||||
- **#153 A12 — la valeur calculée s'affiche sous la formule.** Quand une
|
||||
cellule porte encore le résultat de son dernier calcul Excel, celui-ci
|
||||
s'affiche dans une ligne discrète sous la formule. La seconde lecture
|
||||
`data_only=True` n'a lieu que si l'archive contient réellement une
|
||||
valeur en cache, et toute erreur retombe sur l'affichage formules seul.
|
||||
Info-bulle traduite FR/EN (`xlsx.cached_value_title`).
|
||||
|
||||
---
|
||||
|
||||
## [2.29.0] — 2026-09-27
|
||||
|
||||
### Correction
|
||||
|
||||
+3
-3
@@ -4,7 +4,7 @@
|
||||
|
||||
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
|
||||
|
||||
[]()
|
||||
[]()
|
||||
[](https://opensource.org/licenses/MIT)
|
||||
[](https://www.docker.com/)
|
||||
[](https://www.python.org/)
|
||||
@@ -976,8 +976,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
|
||||
|
||||
## 📝 Changelog
|
||||
|
||||
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.29.0).
|
||||
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.31.0).
|
||||
|
||||
---
|
||||
|
||||
*Projet : ObsiGate | Version : 2.29.0 | Dernière mise à jour : Septembre 2026*
|
||||
*Projet : ObsiGate | Version : 2.31.0 | Dernière mise à jour : Septembre 2026*
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
|
||||
|
||||
[]()
|
||||
[]()
|
||||
[](https://opensource.org/licenses/MIT)
|
||||
[](https://www.docker.com/)
|
||||
[](https://www.python.org/)
|
||||
@@ -1151,8 +1151,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
|
||||
|
||||
## 📝 Changelog
|
||||
|
||||
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.29.0).
|
||||
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.31.0).
|
||||
|
||||
---
|
||||
|
||||
*Project: ObsiGate | Version: 2.29.0 | Last updated: September 2026*
|
||||
*Project: ObsiGate | Version: 2.31.0 | Last updated: September 2026*
|
||||
|
||||
+39
-7
@@ -351,6 +351,23 @@ def _decompress_excalidraw(compressed: str) -> dict[str, Any] | None:
|
||||
return data
|
||||
|
||||
|
||||
def extract_xlsx_indexable(file_path: Path) -> str:
|
||||
"""Return searchable text for a workbook (#153 A5).
|
||||
|
||||
Lazy wrapper: ``openpyxl`` is only imported when a spreadsheet is actually
|
||||
indexed, so a vault without workbooks never pays the import. Errors are
|
||||
swallowed — a corrupt or encrypted file still gets indexed by name.
|
||||
"""
|
||||
try:
|
||||
from backend.xlsx_reader import extract_indexable_text
|
||||
except Exception: # pragma: no cover - openpyxl missing
|
||||
return ""
|
||||
try:
|
||||
return extract_indexable_text(file_path)
|
||||
except Exception: # pragma: no cover - defensive
|
||||
return ""
|
||||
|
||||
|
||||
def extract_excalidraw_indexable(raw: str) -> str:
|
||||
"""Return indexable text content for a raw .excalidraw / .excalidraw.md file.
|
||||
|
||||
@@ -561,11 +578,12 @@ def _scan_vault(
|
||||
title = fpath.stem.replace("-", " ").replace("_", " ")
|
||||
content_preview = ""
|
||||
elif ext == ".xlsx":
|
||||
# #152 — binary workbook: metadata only, the viewer renders
|
||||
# it (parity with _index_single_file_sync).
|
||||
raw = ""
|
||||
# #153 A5 — a workbook stays rendered by the viewer, but its
|
||||
# cell values are now indexed as text so a spreadsheet is
|
||||
# findable by its content (parity with _index_single_file_sync).
|
||||
raw = extract_xlsx_indexable(fpath)
|
||||
title = fpath.stem.replace("-", " ").replace("_", " ")
|
||||
content_preview = ""
|
||||
content_preview = raw[:200].strip()
|
||||
else:
|
||||
raw = fpath.read_text(encoding="utf-8", errors="replace")
|
||||
title = fpath.stem.replace("-", " ").replace("_", " ")
|
||||
@@ -807,6 +825,13 @@ async def reload_index() -> dict[str, Any]:
|
||||
await build_index()
|
||||
# BUG-040/#86: complete the deferred PDF + excalidraw extraction.
|
||||
await enrich_pdf_texts()
|
||||
# The inverted index is NOT updated by the hooks here: the rebuild above
|
||||
# replaces whole vault entries, so the incremental notifications are not
|
||||
# emitted for the files that only changed content. Without this, a manual
|
||||
# reindex left TF-IDF search serving a stale index (BUG-089).
|
||||
from backend.search import init_inverted_index
|
||||
|
||||
init_inverted_index()
|
||||
stats = {}
|
||||
for name, data in index.items():
|
||||
stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])}
|
||||
@@ -882,6 +907,13 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]:
|
||||
# BUG-040/#86: complete the deferred PDF + excalidraw extraction.
|
||||
await enrich_pdf_texts(vault_name)
|
||||
|
||||
# Same as reload_index: the vault entry was replaced wholesale, so rebuild
|
||||
# the inverted index or TF-IDF search keeps serving stale postings
|
||||
# (BUG-089).
|
||||
from backend.search import init_inverted_index
|
||||
|
||||
init_inverted_index()
|
||||
|
||||
stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])}
|
||||
logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags")
|
||||
return stats
|
||||
@@ -955,9 +987,9 @@ def _index_single_file_sync(vault_name: str, vault_path: str, file_path: str, va
|
||||
raw = ""
|
||||
content_preview = ""
|
||||
elif ext == ".xlsx":
|
||||
# #152 — binary workbook: metadata only (parity with _scan_vault).
|
||||
raw = ""
|
||||
content_preview = ""
|
||||
# #153 A5 — index sheet names + header rows as text (see _scan_vault).
|
||||
raw = extract_xlsx_indexable(fpath)
|
||||
content_preview = raw[:200].strip()
|
||||
else:
|
||||
raw = fpath.read_text(encoding="utf-8", errors="replace")
|
||||
content_preview = raw[:200].strip()
|
||||
|
||||
@@ -185,6 +185,26 @@ _ENDPOINT_EXAMPLES: dict[tuple[str, str], dict[str, Any]] = {
|
||||
"request": {"sheet": "Budget", "cells": {"B1": "250"}, "allow_formula": False, "force": False},
|
||||
"response": {"status": "ok", "vault": "TestVault", "path": "data/budget.xlsx", "size": 1},
|
||||
},
|
||||
# GET : pas d'exemple de requête (un requestBody sur un GET serait un OpenAPI
|
||||
# invalide) — les paramètres sont documentés par leurs Query().
|
||||
("get", "/api/file/{vault_name}/xlsx/sheet"): {
|
||||
"response": {
|
||||
"vault": "TestVault",
|
||||
"path": "data/budget.xlsx",
|
||||
"sheet": "Budget",
|
||||
"offset": 0,
|
||||
"limit": 200,
|
||||
"rows": 2,
|
||||
"cols": 2,
|
||||
"total_rows": 640,
|
||||
"total_cols": 12,
|
||||
"max_rows": 500,
|
||||
"max_cols": 40,
|
||||
"truncated": True,
|
||||
"has_more": True,
|
||||
"html": "<table>…</table>",
|
||||
},
|
||||
},
|
||||
("post", "/api/search/replace"): {
|
||||
"request": {"query": "Python", "replacement": "Python 3", "vault": "all", "dry_run": True},
|
||||
"response": {"matches": [{"vault": "TestVault", "path": "note1.md", "title": "Python", "match_count": 3}], "total_matches": 3, "dry_run": True},
|
||||
|
||||
@@ -38,6 +38,7 @@ from backend.schemas import (
|
||||
BrowseResponse,
|
||||
FileContentResponse,
|
||||
FileRawResponse,
|
||||
XlsxSheetWindowResponse,
|
||||
)
|
||||
from backend.services.files import read_raw_file
|
||||
from backend.services.paths import resolve_safe_path
|
||||
@@ -178,6 +179,71 @@ async def api_file_backlinks(
|
||||
}
|
||||
|
||||
|
||||
@router.get(
|
||||
"/api/file/{vault_name}/xlsx/sheet", response_model=XlsxSheetWindowResponse
|
||||
)
|
||||
def api_file_xlsx_sheet(
|
||||
vault_name: str,
|
||||
path: str = Query(..., description="Relative path to the .xlsx file"),
|
||||
sheet: str = Query(..., description="Sheet name (as shown in the viewer tab)"),
|
||||
offset: int = Query(0, ge=0, description="0-based index of the first row to return"),
|
||||
limit: int = Query(
|
||||
200, ge=1, le=1000, description="Rows to return (server-capped)"
|
||||
),
|
||||
current_user=Depends(require_auth),
|
||||
):
|
||||
"""Return a window of rows of one sheet of an .xlsx workbook (#153 A9).
|
||||
|
||||
Backs the viewer's lazy loading: instead of every sheet in a single JSON
|
||||
payload, the client asks for the block it is about to display. The row
|
||||
numbers and the ``data-cell`` references are the real A1 coordinates of the
|
||||
sheet, so a window behaves like the full render (editing a cell in it
|
||||
targets the right cell).
|
||||
|
||||
The response also carries ``total_rows``/``total_cols`` and the ``truncated``
|
||||
flag, so the client can say what is hidden behind the 500x40 render caps
|
||||
instead of silently hiding it.
|
||||
|
||||
Args:
|
||||
vault_name: Name of the vault.
|
||||
path: Relative path of the .xlsx file within the vault.
|
||||
sheet: Sheet name; **404** if the workbook has no such sheet.
|
||||
offset: 0-based index of the first row to return.
|
||||
limit: Rows to return, capped server-side at 1000.
|
||||
|
||||
Returns:
|
||||
``XlsxSheetWindowResponse`` with the rendered ``html`` of the window.
|
||||
|
||||
Raises:
|
||||
HTTPException: 403 (vault access), 404 (vault, file or sheet unknown),
|
||||
415 (not an .xlsx file), 500 (unreadable workbook).
|
||||
"""
|
||||
if not check_vault_access(vault_name, current_user):
|
||||
raise HTTPException(status_code=403, detail=f"Accès refusé à la vault '{vault_name}'")
|
||||
vault_data = get_vault_data(vault_name)
|
||||
if not vault_data:
|
||||
raise HTTPException(status_code=404, detail=f"Vault '{vault_name}' not found")
|
||||
|
||||
file_path = resolve_safe_path(Path(vault_data["path"]), path)
|
||||
if not file_path.is_file():
|
||||
raise HTTPException(status_code=404, detail=f"File not found: {path}")
|
||||
if file_path.suffix.lower() != ".xlsx":
|
||||
raise HTTPException(status_code=415, detail="Le fichier n'est pas un classeur .xlsx")
|
||||
|
||||
# Import tardif : openpyxl n'est chargé que si un .xlsx est réellement demandé.
|
||||
from backend.xlsx_reader import read_sheet_window
|
||||
|
||||
try:
|
||||
window = read_sheet_window(file_path, sheet, offset=offset, limit=limit)
|
||||
except Exception as e:
|
||||
logger.error(f"XLSX sheet read error for {path}: {e}")
|
||||
raise HTTPException(status_code=500, detail=f"Error reading XLSX: {e!s}")
|
||||
if window is None:
|
||||
raise HTTPException(status_code=404, detail=f"Feuille introuvable: {sheet}")
|
||||
|
||||
return {"vault": vault_name, "path": path, **window}
|
||||
|
||||
|
||||
@router.get("/api/file/{vault_name}", response_model=FileContentResponse)
|
||||
async def api_file(vault_name: str, path: str = Query(..., description="Relative path to file"), current_user=Depends(require_auth)):
|
||||
"""Return rendered HTML and metadata for a file.
|
||||
|
||||
+32
-1
@@ -286,7 +286,12 @@ class FileContentResponse(BaseModel):
|
||||
is_csv: bool | None = Field(default=None, description="True for CSV files")
|
||||
is_xlsx: bool | None = Field(default=None, description="True for Excel .xlsx files")
|
||||
xlsx_sheets: list[dict[str, Any]] | None = Field(
|
||||
default=None, description="Rendered xlsx sheets [{name, html}]"
|
||||
default=None,
|
||||
description=(
|
||||
"Rendered xlsx sheets [{name, html, rows, cols, total_rows, "
|
||||
"total_cols, max_rows, max_cols, truncated}] — `truncated` is true "
|
||||
"when the sheet exceeds the 500x40 render caps (#153 A8)"
|
||||
),
|
||||
)
|
||||
xlsx_lossy_features: list[str] | None = Field(
|
||||
default=None,
|
||||
@@ -305,6 +310,32 @@ class FileContentResponse(BaseModel):
|
||||
image_mime: str | None = Field(default=None, description="MIME type for image files")
|
||||
|
||||
|
||||
class XlsxSheetWindowResponse(BaseModel):
|
||||
"""One window of rows of a single .xlsx sheet (lazy loading, #153 A9).
|
||||
|
||||
Served by ``GET /api/file/{vault_name}/xlsx/sheet``; the row numbers and
|
||||
the ``data-cell`` references in ``html`` are the real A1 coordinates of the
|
||||
sheet, whatever the window.
|
||||
"""
|
||||
|
||||
vault: str = Field(description="Vault name")
|
||||
path: str = Field(description="Relative file path within the vault")
|
||||
sheet: str = Field(description="Sheet name (as shown in the tab)")
|
||||
offset: int = Field(description="0-based index of the first returned row")
|
||||
limit: int = Field(description="Maximum number of rows returned (capped server-side)")
|
||||
rows: int = Field(description="Rows actually returned in this window")
|
||||
cols: int = Field(description="Columns of the rendered window")
|
||||
total_rows: int = Field(description="Rows the sheet declares")
|
||||
total_cols: int = Field(description="Columns the sheet declares")
|
||||
max_rows: int = Field(description="Row cap of the renderer (500) — the coverage of this window")
|
||||
max_cols: int = Field(description="Column cap of the renderer (40)")
|
||||
truncated: bool = Field(
|
||||
description="True when the sheet exceeds the 500x40 render caps"
|
||||
)
|
||||
has_more: bool = Field(description="True when rows remain after this window")
|
||||
html: str = Field(description="Rendered HTML table for the window")
|
||||
|
||||
|
||||
class FileRawResponse(BaseModel):
|
||||
"""Raw text content of a file."""
|
||||
|
||||
|
||||
+10
-5
@@ -12,7 +12,12 @@ from sortedcontainers import SortedList
|
||||
|
||||
from backend import indexer as _indexer
|
||||
from backend import semantic_search as _semantic
|
||||
from backend.indexer import index
|
||||
|
||||
# NOTE: the shared index is read through ``_indexer.index`` everywhere, never
|
||||
# via ``from backend.indexer import index``. That import binds the dict object
|
||||
# once, so a module reload of ``backend.indexer`` (tests, dev reload) rebinds
|
||||
# the module-level name to a FRESH dict while this module keeps writing to the
|
||||
# stale one — the inverted index then silently indexes nothing (BUG-089).
|
||||
from backend.services.regex_safety import (
|
||||
MAX_REGEX_MATCHES,
|
||||
truncate_for_regex,
|
||||
@@ -399,7 +404,7 @@ class InvertedIndex:
|
||||
self.vault_docs = defaultdict(set)
|
||||
self.tag_docs = defaultdict(set)
|
||||
|
||||
for vault_name, vault_data in index.items():
|
||||
for vault_name, vault_data in _indexer.index.items():
|
||||
for file_info in vault_data.get("files", []):
|
||||
doc_key = f"{vault_name}::{file_info['path']}"
|
||||
self.doc_count += 1
|
||||
@@ -688,7 +693,7 @@ _indexer.set_index_change_hook(_on_index_change_hook)
|
||||
|
||||
def init_inverted_index():
|
||||
"""Force initial inverted index build. Called after build_index completes on startup."""
|
||||
if any(vdata.get("files") for vdata in index.values()):
|
||||
if any(vdata.get("files") for vdata in _indexer.index.values()):
|
||||
_inverted_index.rebuild()
|
||||
logger.info("Inverted index initialized.")
|
||||
|
||||
@@ -784,7 +789,7 @@ def search(
|
||||
else:
|
||||
candidates = [
|
||||
(vault_name, file_info)
|
||||
for vault_name, vault_data in index.items()
|
||||
for vault_name, vault_data in _indexer.index.items()
|
||||
if vault_filter == "all" or vault_name == vault_filter
|
||||
for file_info in vault_data["files"]
|
||||
]
|
||||
@@ -1613,7 +1618,7 @@ def get_all_tags(vault_filter: str | None = None) -> dict[str, int]:
|
||||
Dict mapping tag names to their total occurrence count.
|
||||
"""
|
||||
merged: dict[str, int] = {}
|
||||
for vault_name, vault_data in index.items():
|
||||
for vault_name, vault_data in _indexer.index.items():
|
||||
if vault_filter and vault_filter != "all" and vault_name != vault_filter:
|
||||
continue
|
||||
for tag, count in vault_data.get("tags", {}).items():
|
||||
|
||||
@@ -19,6 +19,7 @@ import shutil
|
||||
import threading
|
||||
from collections.abc import Callable, Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import date, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
@@ -237,6 +238,14 @@ _XLSX_FLOAT_RE = re.compile(r"^[+-]?(?:\d+\.\d*|\.\d+)$")
|
||||
# the legacy Lotus-style trigger. "+"/"-" are left alone: they are numbers here.
|
||||
_XLSX_FORMULA_RE = re.compile(r"^[=@]")
|
||||
|
||||
# #153 A10 — types recognised when a user types into a cell. Excel infers them
|
||||
# too; storing everything as text would make a spreadsheet unusable (a boolean
|
||||
# column stays a string, a date column sorts lexicographically).
|
||||
_XLSX_TRUE_LITERALS = {"true", "vrai", "oui", "yes"}
|
||||
_XLSX_FALSE_LITERALS = {"false", "faux", "non", "no"}
|
||||
# Shape check before strptime: keeps the hot path free of format attempts.
|
||||
_XLSX_DATE_RE = re.compile(r"^\d{1,2}[-/]\d{1,2}[-/]\d{4}(?:[ T]\d{1,2}:\d{2})?$")
|
||||
|
||||
# #153 A3 — per-file write lock. Two concurrent saves (two tabs, the AI agent
|
||||
# and the viewer, a watcher restore) would otherwise read-modify-write on the
|
||||
# same archive and the last writer silently wins. Kept deliberately small: the
|
||||
@@ -270,7 +279,21 @@ def _xlsx_write_lock(key: str) -> Iterator[None]:
|
||||
|
||||
|
||||
def _coerce_xlsx_value(value: Any) -> Any:
|
||||
"""Turn the string sent by the cell editor back into a scalar."""
|
||||
"""Turn the string sent by the cell editor back into a scalar (#153 A10).
|
||||
|
||||
The coercion is symmetric with :func:`backend.xlsx_reader._fmt`: a value
|
||||
typed by the user comes back as a string, and Excel would have inferred a
|
||||
type when typing the same thing. Recognised here:
|
||||
|
||||
* an empty cell -> ``None`` (clears it)
|
||||
* ``1234`` / ``-1`` -> ``int``
|
||||
* ``1.5`` / ``.5`` -> ``float``
|
||||
* ``TRUE``/``FAUX`` (case-insensitive) -> ``bool``
|
||||
* ``31/12/2026`` / ``31/12/2026 14:30`` -> ``date``/``datetime`` (FR)
|
||||
|
||||
Anything else stays text. A date-looking string typed with a leading
|
||||
``=`` is a formula and never reaches here as a date.
|
||||
"""
|
||||
if not isinstance(value, str):
|
||||
return value
|
||||
text = value.strip()
|
||||
@@ -280,9 +303,35 @@ def _coerce_xlsx_value(value: Any) -> Any:
|
||||
return int(text)
|
||||
if _XLSX_FLOAT_RE.match(text):
|
||||
return float(text)
|
||||
lowered = text.lower()
|
||||
if lowered in _XLSX_TRUE_LITERALS:
|
||||
return True
|
||||
if lowered in _XLSX_FALSE_LITERALS:
|
||||
return False
|
||||
if not _XLSX_FORMULA_RE.match(text):
|
||||
parsed = _parse_fr_datetime(text)
|
||||
if parsed is not None:
|
||||
return parsed
|
||||
return value
|
||||
|
||||
|
||||
def _parse_fr_datetime(text: str) -> date | datetime | None:
|
||||
"""Parse a FR-localised date/datetime, or return ``None``.
|
||||
|
||||
Accepts ``JJ/MM/AAAA`` and ``JJ/MM/AAAA HH:MM`` (also ``JJ-MM-AAAA``).
|
||||
``dayfirst`` is what makes ``01/02/2026`` the 1st of February rather than
|
||||
the 2nd of January — the French convention.
|
||||
"""
|
||||
if not _XLSX_DATE_RE.match(text):
|
||||
return None
|
||||
for fmt in ("%d/%m/%Y %H:%M", "%d/%m/%Y", "%d-%m-%Y %H:%M", "%d-%m-%Y"):
|
||||
try:
|
||||
return datetime.strptime(text, fmt)
|
||||
except ValueError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _write_cell(ws: Any, ref: str, value: Any, *, allow_formula: bool) -> None:
|
||||
"""Assign one cell, forcing text when it looks like a formula.
|
||||
|
||||
|
||||
+302
-15
@@ -11,6 +11,7 @@ round-trip would drop (#153 A1) so the UI can warn before saving.
|
||||
from __future__ import annotations
|
||||
|
||||
import html
|
||||
import logging
|
||||
import re
|
||||
import zipfile
|
||||
from datetime import date, datetime
|
||||
@@ -20,11 +21,19 @@ from typing import Any
|
||||
from openpyxl import load_workbook
|
||||
from openpyxl.utils import get_column_letter
|
||||
|
||||
logger = logging.getLogger("obsigate.xlsx_reader")
|
||||
|
||||
# ponytail: hard caps bound the rendered grid (500 rows x 40 cols per sheet).
|
||||
# Raise them, or paginate per sheet, if a real workbook needs more.
|
||||
MAX_ROWS = 500
|
||||
MAX_COLS = 40
|
||||
|
||||
# #153 A9 — window size served by ``read_sheet_window()`` (lazy per-sheet
|
||||
# loading). The endpoint is bounded so a single request can never ask for the
|
||||
# whole workbook back in one JSON payload; the UI pages through the rest.
|
||||
MAX_WINDOW_ROWS = 1_000
|
||||
DEFAULT_WINDOW_ROWS = 200
|
||||
|
||||
# #153 A1 — workbook parts openpyxl does not re-serialize on load+save.
|
||||
# Verified against openpyxl 3.1.5: charts, images, drawings and pivot tables
|
||||
# DO survive the round-trip, so they are deliberately absent from this map.
|
||||
@@ -48,6 +57,14 @@ _CACHED_FORMULA_RE = re.compile(rb"<f[ >][^<]*</f>\s*<v>[^<]")
|
||||
# Sheet XML scanned by the cached-formula probe (CPU guard, like MAX_REPLACE_FILE_BYTES).
|
||||
_MAX_PROBE_BYTES = 8_000_000
|
||||
|
||||
# #153 A5 — ceiling on the text handed to the TF-IDF / semantic index. A workbook
|
||||
# is a data dump, not prose: indexing every cell would flood the inverted index
|
||||
# and bury the notes. Sheet names + the first rows are enough to make a
|
||||
# spreadsheet findable by its headers.
|
||||
MAX_INDEX_CHARS = 5_000
|
||||
_INDEX_ROWS_PER_SHEET = 20
|
||||
MAX_INDEX_SHEETS = 20
|
||||
|
||||
|
||||
def _fmt(value: Any) -> str:
|
||||
if value is None:
|
||||
@@ -74,7 +91,37 @@ def _trim(grid: list[list[str]]) -> list[list[str]]:
|
||||
return [row[:width] for row in grid]
|
||||
|
||||
|
||||
def _table(grid: list[list[str]]) -> str:
|
||||
def _cell_cached(cached: list[list[str]] | None, r: int, c: int) -> str:
|
||||
"""Return the cached result for a 0-based cell, or ``""``.
|
||||
|
||||
The shadow grid is read positionally and may be narrower than the formula
|
||||
grid (``_trim`` collapses the trailing empty columns of each grid
|
||||
independently), so every lookup is bounds-checked rather than assumed.
|
||||
"""
|
||||
if not cached or r >= len(cached):
|
||||
return ""
|
||||
row = cached[r]
|
||||
return row[c] if c < len(row) else ""
|
||||
|
||||
|
||||
def _table(
|
||||
grid: list[list[str]],
|
||||
cached: list[list[str]] | None = None,
|
||||
row_offset: int = 0,
|
||||
) -> str:
|
||||
"""Render a grid as an HTML table.
|
||||
|
||||
``cached`` is the same grid read with ``data_only=True`` (#153 A12): where a
|
||||
formula cell still carries its last computed result, it is shown as a
|
||||
discreet second line (``<span class="xlsx-cached">``) so the user sees the
|
||||
number Excel last calculated instead of only the formula text. The span
|
||||
carries ``data-cached-value`` and is titled client-side from
|
||||
``xlsx.cached_value_title`` — the backend never emits UI text.
|
||||
|
||||
``row_offset`` is the number of rows skipped before this grid (#153 A9): the
|
||||
row numbers and the ``data-cell`` references must stay the real A1
|
||||
coordinates of the sheet, not of the window.
|
||||
"""
|
||||
if not grid:
|
||||
return "<p><em>Feuille vide</em></p>"
|
||||
n_cols = max(len(row) for row in grid)
|
||||
@@ -86,11 +133,28 @@ def _table(grid: list[list[str]]) -> str:
|
||||
]
|
||||
out += [f"<th>{get_column_letter(c)}</th>" for c in range(1, n_cols + 1)]
|
||||
out.append("</tr></thead><tbody>")
|
||||
for r, row in enumerate(grid, start=1):
|
||||
for r, row in enumerate(grid, start=row_offset + 1):
|
||||
out.append(f'<tr><th class="xlsx-rownum">{r}</th>')
|
||||
for c, val in enumerate(row, start=1):
|
||||
ref = f"{get_column_letter(c)}{r}"
|
||||
out.append(f'<td data-cell="{ref}">{html.escape(val)}</td>')
|
||||
# The cached result only makes sense for a formula cell: on a plain
|
||||
# value cell the two reads are identical and showing both would
|
||||
# duplicate the text.
|
||||
shadow = ""
|
||||
if cached is not None and val.startswith("="):
|
||||
# `c` is 1-based (A1 notation) and `r` too, while the grid is
|
||||
# 0-based: translate both.
|
||||
cval = _cell_cached(cached, r - 1, c - 1)
|
||||
if cval and cval != val:
|
||||
# The tooltip is translated client-side from
|
||||
# `xlsx.cached_value_title`; never hardcode UI text here.
|
||||
shadow = (
|
||||
f'<span class="xlsx-cached" data-cached-value="1">'
|
||||
f"{html.escape(cval)}</span>"
|
||||
)
|
||||
out.append(
|
||||
f'<td data-cell="{ref}">{html.escape(val)}{shadow}</td>'
|
||||
)
|
||||
out.append("</tr>")
|
||||
out.append("</tbody></table></div>")
|
||||
return "".join(out)
|
||||
@@ -141,19 +205,242 @@ def inspect_workbook(file_path: Path) -> list[str]:
|
||||
return []
|
||||
|
||||
|
||||
def render_sheets(file_path: Path) -> list[dict[str, str]]:
|
||||
"""Return ``[{"name": sheet_title, "html": table_html}, ...]``."""
|
||||
def render_sheets(file_path: Path) -> list[dict[str, Any]]:
|
||||
"""Return one dict per sheet: ``{name, html, rows, cols, total_*, truncated}``.
|
||||
|
||||
Reads the workbook twice: once with ``data_only=False`` for the formulas
|
||||
(what the user must edit) and, when any formula carries a cached result
|
||||
(#153 A12), once with ``data_only=True`` to show what Excel last computed.
|
||||
The second pass is skipped entirely when the archive holds no cached value,
|
||||
so the common case still costs a single load.
|
||||
|
||||
``total_rows``/``total_cols`` are the dimensions the sheet declares and
|
||||
``truncated`` says whether the hard caps actually cut it (#153 A8) — the
|
||||
viewer needs both to stop silently hiding the tail of a sheet.
|
||||
"""
|
||||
wb = load_workbook(str(file_path), read_only=True, data_only=False)
|
||||
try:
|
||||
sheets = []
|
||||
for ws in wb.worksheets:
|
||||
grid = [
|
||||
[_fmt(v) for v in row]
|
||||
for row in ws.iter_rows(
|
||||
min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True
|
||||
)
|
||||
]
|
||||
sheets.append({"name": ws.title, "html": _table(_trim(grid))})
|
||||
return sheets
|
||||
formulas = [_sheet_grid(ws) for ws in wb.worksheets]
|
||||
titles = [ws.title for ws in wb.worksheets]
|
||||
extents = [_sheet_extent(ws) for ws in wb.worksheets]
|
||||
finally:
|
||||
wb.close()
|
||||
|
||||
cached: list[list[list[str]]] | None = None
|
||||
if _has_cached_values(file_path):
|
||||
cached = _read_cached_grids(file_path, titles)
|
||||
|
||||
sheets = []
|
||||
for i, title in enumerate(titles):
|
||||
grid = _trim(formulas[i])
|
||||
# The shadow grid is NOT trimmed independently: _trim drops the
|
||||
# trailing empty columns of each grid on its own width, which would
|
||||
# shift every cached value left of its formula. Indexing it
|
||||
# positionally against the untrimmed grid keeps the two aligned.
|
||||
shadow = cached[i] if cached is not None and i < len(cached) else None
|
||||
total_rows, total_cols = extents[i]
|
||||
sheets.append(
|
||||
{
|
||||
"name": title,
|
||||
"html": _table(grid, shadow),
|
||||
"rows": len(grid),
|
||||
"cols": max((len(r) for r in grid), default=0),
|
||||
"total_rows": total_rows,
|
||||
"total_cols": total_cols,
|
||||
# Coverage, not display size: `rows`/`cols` are post-trim (a
|
||||
# sheet of 3 filled cells in a 500-row block renders 1x1), and
|
||||
# the client must announce the cap it stopped at, not how many
|
||||
# cells happen to be non-empty.
|
||||
"max_rows": MAX_ROWS,
|
||||
"max_cols": MAX_COLS,
|
||||
# A sheet is truncated when the caps, not the trailing blanks,
|
||||
# decided its shape: comparing against the *rendered* size would
|
||||
# flag every sheet carrying a few empty formatted rows.
|
||||
"truncated": total_rows > MAX_ROWS or total_cols > MAX_COLS,
|
||||
}
|
||||
)
|
||||
return sheets
|
||||
|
||||
|
||||
def read_sheet_window(
|
||||
file_path: Path,
|
||||
sheet: str,
|
||||
offset: int = 0,
|
||||
limit: int = DEFAULT_WINDOW_ROWS,
|
||||
) -> dict[str, Any] | None:
|
||||
"""Return a window of rows of one sheet, or ``None`` if the sheet is unknown.
|
||||
|
||||
Backs the lazy per-sheet loading of #153 A9: the viewer asks for the rows
|
||||
it is about to display instead of shipping every sheet in the initial file
|
||||
payload. ``offset`` is 0-based; the row numbers and the ``data-cell``
|
||||
references in the returned ``html`` are the real A1 coordinates of the
|
||||
sheet, so a window is indistinguishable from a full render.
|
||||
|
||||
``limit`` is clamped to :data:`MAX_WINDOW_ROWS`. Raises nothing: an unknown
|
||||
sheet yields ``None`` and a broken workbook propagates the caller's usual
|
||||
500.
|
||||
"""
|
||||
offset = max(int(offset), 0)
|
||||
limit = min(max(int(limit), 1), MAX_WINDOW_ROWS)
|
||||
|
||||
wb = load_workbook(str(file_path), read_only=True, data_only=False)
|
||||
try:
|
||||
if sheet not in wb.sheetnames:
|
||||
return None
|
||||
ws = wb[sheet]
|
||||
total_rows, total_cols = _sheet_extent(ws)
|
||||
grid = _trim(
|
||||
_sheet_grid(ws, min_row=offset + 1, max_row=offset + limit)
|
||||
)
|
||||
finally:
|
||||
wb.close()
|
||||
|
||||
shadow: list[list[str]] | None = None
|
||||
# Same A12 rule as the full render: the second read only happens when the
|
||||
# archive really holds cached results.
|
||||
if _has_cached_values(file_path):
|
||||
shadow = _read_cached_window(file_path, sheet, offset, limit)
|
||||
return {
|
||||
"sheet": sheet,
|
||||
"offset": offset,
|
||||
"limit": limit,
|
||||
"rows": len(grid),
|
||||
"cols": max((len(r) for r in grid), default=0),
|
||||
"total_rows": total_rows,
|
||||
"total_cols": total_cols,
|
||||
"max_rows": MAX_ROWS,
|
||||
"max_cols": MAX_COLS,
|
||||
"truncated": total_rows > MAX_ROWS or total_cols > MAX_COLS,
|
||||
"has_more": offset + len(grid) < total_rows,
|
||||
"html": _table(grid, shadow, row_offset=offset),
|
||||
}
|
||||
|
||||
|
||||
def _read_cached_window(
|
||||
file_path: Path, sheet: str, offset: int, limit: int
|
||||
) -> list[list[str]] | None:
|
||||
"""``data_only=True`` grid for one window, or ``None`` if unavailable.
|
||||
|
||||
Best effort like :func:`_read_cached_grids`: a workbook Excel opens but
|
||||
openpyxl cannot re-read must still display (formulas only).
|
||||
"""
|
||||
try:
|
||||
wb = load_workbook(str(file_path), read_only=True, data_only=True)
|
||||
except Exception:
|
||||
return None
|
||||
try:
|
||||
if sheet not in wb.sheetnames:
|
||||
return None
|
||||
return _sheet_grid(
|
||||
wb[sheet], min_row=offset + 1, max_row=offset + limit
|
||||
)
|
||||
except Exception:
|
||||
logger.debug("xlsx cached window unavailable", exc_info=True)
|
||||
return None
|
||||
finally:
|
||||
wb.close()
|
||||
|
||||
|
||||
def _sheet_extent(ws: Any) -> tuple[int, int]:
|
||||
"""Rows and columns the worksheet declares, never negative.
|
||||
|
||||
``max_row``/``max_column`` come from the sheet's dimension record; a
|
||||
hand-edited file may omit it, hence the defensive coercion.
|
||||
"""
|
||||
try:
|
||||
rows = max(int(getattr(ws, "max_row", 0) or 0), 0)
|
||||
except (TypeError, ValueError):
|
||||
rows = 0
|
||||
try:
|
||||
cols = max(int(getattr(ws, "max_column", 0) or 0), 0)
|
||||
except (TypeError, ValueError):
|
||||
cols = 0
|
||||
return rows, cols
|
||||
|
||||
|
||||
def _sheet_grid(
|
||||
ws: Any, min_row: int = 1, max_row: int = MAX_ROWS, max_col: int = MAX_COLS
|
||||
) -> list[list[str]]:
|
||||
"""Read a worksheet window into a grid of formatted strings, bounded by the caps."""
|
||||
return [
|
||||
[_fmt(v) for v in row]
|
||||
for row in ws.iter_rows(
|
||||
min_row=min_row, max_row=max_row, max_col=max_col, values_only=True
|
||||
)
|
||||
]
|
||||
|
||||
|
||||
def _has_cached_values(file_path: Path) -> bool:
|
||||
"""True when the archive holds at least one ``<f>…</f><v>…</v>``."""
|
||||
try:
|
||||
with zipfile.ZipFile(file_path) as zf:
|
||||
return _has_cached_formulas(zf)
|
||||
except (OSError, zipfile.BadZipFile):
|
||||
return False
|
||||
|
||||
|
||||
def _read_cached_grids(
|
||||
file_path: Path, titles: list[str]
|
||||
) -> list[list[list[str]]] | None:
|
||||
"""Read every sheet with ``data_only=True`` (what Excel last computed).
|
||||
|
||||
Best effort: returns ``None`` on any failure so the viewer falls back to the
|
||||
formula-only rendering. A workbook Excel opens but openpyxl cannot re-read
|
||||
must still display.
|
||||
"""
|
||||
try:
|
||||
wb = load_workbook(str(file_path), read_only=True, data_only=True)
|
||||
except Exception:
|
||||
return None
|
||||
try:
|
||||
grids = [_sheet_grid(ws) for ws in wb.worksheets]
|
||||
if [ws.title for ws in wb.worksheets] != titles:
|
||||
return None
|
||||
return grids
|
||||
except Exception:
|
||||
logger.debug("xlsx cached values unavailable", exc_info=True)
|
||||
return None
|
||||
finally:
|
||||
wb.close()
|
||||
|
||||
|
||||
def extract_indexable_text(file_path: Path) -> str:
|
||||
"""Return searchable text for the TF-IDF / semantic index (#153 A5).
|
||||
|
||||
Sheet names plus the first :data:`_INDEX_ROWS_PER_SHEET` rows of each
|
||||
sheet, capped at :data:`MAX_INDEX_CHARS`. Rows are tab-joined so a search
|
||||
for a header matches the sheet it belongs to.
|
||||
|
||||
Never raises: a corrupt, encrypted or unsupported workbook yields ``""`` so
|
||||
the file still gets indexed by name (same contract as :func:`inspect_workbook`).
|
||||
"""
|
||||
chunks: list[str] = []
|
||||
budget = MAX_INDEX_CHARS
|
||||
try:
|
||||
wb = load_workbook(str(file_path), read_only=True, data_only=True)
|
||||
except Exception:
|
||||
# Encrypted (BadZipFile) or not a real workbook: name-only indexing.
|
||||
return ""
|
||||
try:
|
||||
for ws in wb.worksheets[:MAX_INDEX_SHEETS]:
|
||||
if budget <= 0:
|
||||
break
|
||||
# The sheet title alone is a strong signal ("Recettes", "Budget").
|
||||
block = [ws.title]
|
||||
for row in ws.iter_rows(
|
||||
min_row=1, max_row=_INDEX_ROWS_PER_SHEET, max_col=MAX_COLS, values_only=True
|
||||
):
|
||||
cells = [_fmt(v) for v in row]
|
||||
# Skip blank rows instead of emitting runs of tabs.
|
||||
if not any(c.strip() for c in cells):
|
||||
continue
|
||||
block.append("\t".join(cells).rstrip())
|
||||
text = "\n".join(block)
|
||||
chunks.append(text[:budget])
|
||||
budget -= len(text)
|
||||
except Exception:
|
||||
# Truncated but still useful: keep whatever was collected.
|
||||
pass
|
||||
finally:
|
||||
wb.close()
|
||||
return "\n".join(c for c in chunks if c).strip()
|
||||
|
||||
Generated
+1
-1
@@ -2626,7 +2626,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "obsigate-desktop"
|
||||
version = "2.29.0"
|
||||
version = "2.31.0"
|
||||
dependencies = [
|
||||
"chrono",
|
||||
"env_logger",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "obsigate-desktop"
|
||||
version = "2.29.0"
|
||||
version = "2.31.0"
|
||||
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
|
||||
authors = ["Bruno Charest"]
|
||||
edition = "2021"
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
|
||||
"productName": "ObsiGate",
|
||||
"version": "2.29.0",
|
||||
"version": "2.31.0",
|
||||
"identifier": "com.obsigate.desktop",
|
||||
"build": {
|
||||
"frontendDist": "../frontend",
|
||||
|
||||
@@ -171,13 +171,36 @@ curl -X PUT "http://localhost:2020/api/file/Recettes/xlsx/save?path=budget.xlsx"
|
||||
- Deux sauvegardes simultanées sur le même fichier : la seconde reçoit
|
||||
**409** `conflict` au lieu d'écraser la première.
|
||||
|
||||
### Feuilles volumineuses et lecture par fenêtres
|
||||
|
||||
Le rendu est plafonné à **500 lignes × 40 colonnes** par feuille. Quand
|
||||
une feuille dépasse ce plafond, un bandeau **« Feuille tronquée »**
|
||||
l'annonce explicitement (par exemple « 500 lignes affichées sur 520 »)
|
||||
au lieu de présenter une table courte comme complète — le classeur,
|
||||
lui, n'est jamais modifié. La ligne d'en-têtes de colonnes reste
|
||||
visible pendant le défilement vertical.
|
||||
|
||||
Côté API, `GET /api/file/{vault}/xlsx/sheet` sert une feuille **par
|
||||
fenêtres de lignes**, y compris au-delà du plafond d'affichage — les
|
||||
coordonnées A1 renvoyées sont celles de la feuille réelle :
|
||||
|
||||
```bash
|
||||
curl "http://localhost:2020/api/file/Recettes/xlsx/sheet?path=budget.xlsx&sheet=Budget&offset=500&limit=200"
|
||||
```
|
||||
|
||||
- `offset` : première ligne renvoyée (0-based) ; `limit` : nombre de
|
||||
lignes (1 à 1 000 par requête).
|
||||
- La réponse porte `total_rows`, `truncated` et `has_more` pour paginer.
|
||||
- Erreurs : **404** si la feuille n'existe pas, **415** si le fichier
|
||||
n'est pas un `.xlsx`.
|
||||
|
||||
### Limites
|
||||
|
||||
- Le rendu est plafonné à **500 lignes × 40 colonnes** par feuille, sans
|
||||
pagination : au-delà, le contenu n'est pas affiché (et non éditable).
|
||||
- L'affichage intégré reste plafonné à **500 lignes × 40 colonnes** par
|
||||
feuille (le défilement automatique au-delà est en préparation) ; les
|
||||
lignes cachées restent accessibles via l'endpoint ci-dessus.
|
||||
- Styles, formats de nombre, cellules fusionnées et volets figés ne sont pas
|
||||
rendus ; le contenu des tableurs n'est pas non plus indexé pour la
|
||||
recherche (contrairement aux PDF).
|
||||
rendus.
|
||||
- Formats non gérés : `.xls`, `.xlsm` (macros), `.ods`.
|
||||
|
||||
---
|
||||
|
||||
@@ -195,6 +195,8 @@ Avant de corriger quoi que ce soit, un agent IA doit :
|
||||
| *BUG-086* | Édition d'un `.xlsx` : `wb.save()` écrit en place, un plantage laisse un classeur corrompu | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::edit_xlsx_cells` | Simuler un `OSError` pendant `Workbook.save` → le fichier d'origine est tronqué | Écriture atomique : `wb.save(<nom>.<pid>.tmp)` puis `os.replace()` ; `.tmp` supprimé sur échec ; le backup `.bak` reste inchangé. Test : `TestXlsxAtomicWrite::test_failed_save_keeps_the_original` (octets identiques après échec) + `test_no_tmp_left_after_a_successful_save` | #153 A2. Le fichier temporaire a un suffixe `.tmp` → ignoré par le watcher (`_is_relevant` ne retient que les extensions supportées). Vérifié : cf. BUG-085 |
|
||||
| *BUG-087* | Édition d'un `.xlsx` concurrente (deux onglets, agent IA + viewer) : read-modify-write sans verrou, le dernier écrivain gagne silencieusement | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::_xlsx_write_lock` | Deux `PUT xlsx/save` simultanés sur le même fichier → une écriture est écrasée sans trace | Verrou par chemin (registre + garde, timeout 15 s) autour du cycle load → edit → `os.replace` ; attente dépassée → **409** `conflict`. L'endpoint est devenu `def` (sync) pour que l'attente s'exécute dans le threadpool et ne bloque pas la boucle d'événements. Test : `TestXlsxWriteLock` (2) | #153 A3. Verrou en mémoire, par processus : protège les cas d'un même serveur (le cas desktop/Tauri). Vérifié : cf. BUG-085 |
|
||||
| *BUG-088* | Injection de formule dans un `.xlsx` : une saisie `=cmd\|'/c calc'!A1` est stockée comme formule et s'exécute à l'ouverture dans Excel (DDE) | 🟢 corrigé | P0 | tableur Excel / sécurité | IA | `backend/services/mutations.py::_write_cell`, `backend/routers/files_write.py`, `frontend/js/viewer.js::renderXlsxViewer` | `PUT /api/file/V/xlsx/save` avec `{"sheet": "S", "cells": {"A1": "=1+1"}}` → la cellule sort en `data_type == "f"` | `cell.data_type = "s"` après affectation : le texte est stocké comme chaîne, aucun `<f>` n'est écrit. Opt-in via `allow_formula: true` (endpoint) et le bouton `f(x)` de la visionneuse (session, jamais persisté). Test : `TestXlsxFormulaGuard` (4) + `xlsx-viewer.test.mjs` (toggle) | #153 A4. `+`/`-` ne sont pas neutralisés : ils sont déjà convertis en nombre par `_coerce_xlsx_value`. Le handler global `ServiceError` expose désormais `code` + `details` (le client en a besoin pour le 409), et `api()` (frontend) les propage sur l'Error. Vérifié : cf. BUG-085 |
|
||||
| *BUG-089* | Un reindex manuel ne reconstruisait pas l'index inversé : la recherche TF-IDF continuait de servir un index périmé | 🟢 corrigé | P1 | ⚙️ backend / recherche | IA | `backend/indexer.py::reload_index`, `backend/indexer.py::reload_single_vault`, `backend/search.py` | Modifier le contenu d'un fichier, puis `GET /api/index/reload` → la recherche renvoie encore l'ancien contenu (ou rien pour un fichier nouveau) | `reload_index()` / `reload_single_vault()` appellent `init_inverted_index()` après le rebuild (le remplacement wholesale d'une entrée de vault n'émet pas les notifications incrémentales). En prime, `backend/search.py` lisait l'index via `from backend.indexer import index` (liaison **par valeur** du dict) : un `importlib.reload(backend.indexer)` recréait le dict côté indexer tandis que la recherche écrivait encore dans l'ancien — l'index inversé n'indexait alors plus rien. Tous les accès passent désormais par `_indexer.index`. Contre-preuve : `TestXlsxSearchable::test_search_finds_a_word_stored_in_a_cell` échoue sans le correctif | #153 A5. Trouvé en écrivant le test de recherche d'A5 : il passait isolément et échouait en suite complète selon l'ordre. Le reload incrémental par fichier (watcher, edition) n'est pas concerné : il passe par le hook `_on_index_change`. Vérifié : suite 1402 passed / 6 skipped, ruff/mypy 0 |
|
||||
| *BUG-090* | Troncature silencieuse d'une feuille `.xlsx` au-delà de 500 lignes × 40 colonnes : l'utilisateur voit une table courte sans aucun indice que la suite existe | 🟢 corrigé | P1 | tableur Excel / UX | IA | `backend/xlsx_reader.py::render_sheets`, `backend/routers/files_read.py`, `frontend/js/viewer.js::renderXlsxViewer`, `frontend/style.css` | Ouvrir `test_vault/sample-xlsx-large.xlsx` (520 lignes) → la feuille s'arrête à la ligne 500 sans aucun message | `render_sheets()` renvoie désormais `total_rows`/`total_cols` (dimensions déclarées par la feuille), `max_rows`/`max_cols` (plafonds du moteur) et `truncated` ; la visionneuse affiche un bandeau « Feuille tronquée — 500 lignes affichées sur 520 » (i18n `xlsx.truncated_*` FR/EN, axe des colonnes inclus). Contre-preuve : neutraliser `truncated` → `TestXlsxTruncationNotice` (2 tests) échoue | #153 A8/R5. La ligne d'en-têtes est aussi `sticky` au défilement vertical (`thead th { top: 0 }` + `top: auto` sur les numéros de ligne pour éviter l'empilement en haut à gauche). L'endpoint `GET …/xlsx/sheet` (#153 A9) sert les fenêtres au-delà du plafond, mais le chargement paresseux complet (défilement virtuel, « charger tout ») reste à faire — le bandeau dit la vérité en attendant. Vérifié : `test_xlsx_viewer.py` 58 passed, E2E 7/7 (dont 3 nouveaux), suite 1417 passed / 6 skipped, ruff/mypy 0, i18n parity |
|
||||
|
||||
### TODOs techniques (améliorations / nouvelles tâches)
|
||||
|
||||
@@ -212,6 +214,8 @@ Avant de corriger quoi que ce soit, un agent IA doit :
|
||||
|
||||
| Date | ID(s) traité(s) | Action | Fichiers modifiés | Résumé | Statut après |
|
||||
|---|---|---|---|---|---|
|
||||
| 2026-09-28 | BUG-090 (#153 A8 + A9) | Correction + feature | `backend/xlsx_reader.py`, `backend/routers/files_read.py`, `backend/schemas.py`, `backend/openapi_docs.py`, `frontend/js/viewer.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-large.xlsx` | **La troncature d'une feuille est annoncée et les lignes cachées restent accessibles** : (BUG-090/A8) `render_sheets()` renvoie `total_rows`/`total_cols`/`max_rows`/`max_cols`/`truncated`, la visionneuse affiche un bandeau « Feuille tronquée » (i18n FR/EN, axes lignes et colonnes) et la ligne d'en-têtes devient `sticky` (`top: auto` sur les numéros de ligne pour éviter l'empilement) ; (A9) `GET /api/file/{vault}/xlsx/sheet?sheet=&offset=&limit=` (`XlsxSheetWindowResponse`, plafond 1 000 lignes/requête, 404 feuille inconnue, 415 non-xlsx) sert une fenêtre avec les **vraies** coordonnées A1 et le `has_more` de pagination. Contre-preuves : neutraliser `truncated` → 2 tests échouent ; neutraliser l'offset → 3 tests échouent. Vérifié : `test_xlsx_viewer.py` 58 passed, xlsx-viewer.test.mjs 14/14, E2E 7/7 (3 nouveaux + fixture `sample-xlsx-large.xlsx` 520 lignes), suite 1417 passed / 6 skipped, ruff 0, mypy 0, i18n parity, validate-imports 40 modules | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-28 | BUG-089 (#153 A5, A10, A12) | Correction | `backend/xlsx_reader.py`, `backend/indexer.py`, `backend/search.py`, `backend/services/mutations.py`, `frontend/js/viewer.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/test_xlsx_viewer.py` | **Les tableurs deviennent visibles ettypés** : (A5) `extract_indexable_text()` indexe noms de feuilles + 20 premières lignes (plafond 5 k caractères) dans le TF-IDF et la recherche sémantique — un mot tapé dans une cellule rend le fichier trouvable ; (A10) `_coerce_xlsx_value()` reconnaît désormais les booléens (`TRUE`/`FAUX`/`OUI`/`NON`) et les dates FR `JJ/MM/AAAA` (jour-first : `01/02/2026` = 1er février), symétrique avec l'affichage ; (A12) la valeur calculée en cache s'affiche sous la formule (`<span class="xlsx-cached">`, 2ᵉ lecture `data_only=True` uniquement si l'archive contient un `<v>`), info-bulle traduite via `xlsx.cached_value_title` FR/EN. (BUG-089) un reindex manuel reconstruisait mal l'index inversé et `backend/search.py` lisait l'index par valeur. Contre-preuves vérifiées pour A5, A10 et A12. Vérifié : `test_xlsx_viewer.py` 43 passed, suite 1402 passed / 6 skipped, ruff 0, mypy 0, i18n parity, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10 | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| 2026-09-27 | BUG-085 → BUG-088 (#153 A1-A4) | Correction | `backend/xlsx_reader.py`, `backend/services/mutations.py`, `backend/routers/files_read.py`, `backend/routers/files_write.py`, `backend/schemas.py`, `backend/main.py`, `frontend/js/viewer.js`, `frontend/js/auth.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `frontend/sw.js`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-lossy.xlsx`, `.gitea/workflows/ci.yml` | **Garde-fous d'écriture des classeurs Excel** : (BUG-085) `inspect_workbook()` détecte ce qu'un round-trip openpyxl perd (valeurs calculées, slicers, contrôles, connexions, custom XML, signature) → la lecture expose `xlsx_lossy_features`, la visionneuse affiche une bannière et `PUT xlsx/save` refuse sans `force` (**409** `xlsx_lossy_content`, confirmation explicite puis reprise) ; (BUG-086) écriture atomique `.tmp` + `os.replace` ; (BUG-087) verrou par fichier (409 `conflict`, endpoint sync pour le threadpool) ; (BUG-088) une saisie `=`/`@` est stockée en texte (`data_type = "s"`), sauf opt-in `allow_formula` / bouton `f(x)`. Le handler `ServiceError` expose désormais `code` + `details` et `api()` les propage. Périmètre de perte revalidé empiriquement sur openpyxl 3.1.5 (graphiques, images et TCD sont préservés). Vérifié : `test_xlsx_viewer.py` 31 passed, suite 1390 passed / 6 skipped, ruff/mypy 0, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10, E2E 3/3 | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
| *(exemple)* 2026-06-15 | BUG-001 | Correction | `frontend/app.js` | Réécriture de `renderFile()` pour préserver le DOM dashboard | 🟢 corrigé (en attente vérif) |
|
||||
| 2026-09-09 | BUG-001, BUG-002 | Correction | `backend/main.py`, `frontend/excalidraw-editor.html`, `tests/test_pdf_stream.py` | BUG-001: Content-Disposition RFC 5987 (nom PDF accentué ne casse plus l'en-tête → plus de 500). BUG-002: suppression alias esm.sh (408 jotai) + React 19 cohérent + prop `excalidrawAPI` → Loading masqué, save OK. Vérifié: 534 tests backend verts + E2E navigateur. | 🟢 corrigé (en attente vérif utilisateur) |
|
||||
|
||||
+9
-9
@@ -1,6 +1,6 @@
|
||||
# ObsiGate — Roadmap
|
||||
|
||||
> **Version :** 2.29.0 | **Dernière mise à jour :** 2026-09-27
|
||||
> **Version :** 2.31.0 | **Dernière mise à jour :** 2026-09-28
|
||||
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
|
||||
> vers les fonctionnalités livrées.
|
||||
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
|
||||
@@ -47,7 +47,7 @@
|
||||
### 153. Visionneuse & édition XLSX — complétude (fidélité, recherche, IA, UX, formats)
|
||||
|
||||
- **Effort :** 8-13 jours (P0 ✅ 2-3 j · P1 : 4-6 j · P2 : 2-4 j) | **Impact :** 🟡
|
||||
- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), reste P1 puis P2
|
||||
- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), **A8/A9 livrés le 2026-09-28** (avec BUG-090, défilement virtuel A9bis à venir), reste A6-A7 puis A13-A17
|
||||
- **Analyse, risques et critères d'acceptation :** [features/xlsx-viewer.md](./features/xlsx-viewer.md)
|
||||
- **Description :** #152 (visionneuse XLSX, 2.27.0) lit et édite correctement la **grille de
|
||||
valeurs** d'un `.xlsx`, mais l'ensemble supporté est étroit : valeurs seulement (ni structure,
|
||||
@@ -70,15 +70,15 @@
|
||||
- [x] **A2** Écriture atomique (`wb.save(.tmp)` + `os.replace()`, backup inchangé) — BUG-086
|
||||
- [x] **A3** Verrou par fichier autour du read-modify-write (timeout 15 s + **409** `conflict`) — BUG-087
|
||||
- [x] **A4** Neutralisation de l'injection de formule (`=`/`@` stockés en texte, opt-in `allow_formula` + bouton `f(x)`) — BUG-088
|
||||
- **P1 — recherche, IA, UX (🟡, 4-6 j) — ⚪ à faire**
|
||||
- [ ] **A5** Indexation du contenu des feuilles (TF-IDF + sémantique, plafond ~5 k caractères)
|
||||
- **P1 — recherche, IA, UX (🟡, 4-6 j) — 🔵 en cours**
|
||||
- [x] **A5** Indexation du contenu des feuilles (noms de feuilles + 20 premières lignes, plafond 5 k caractères) — les mots tapés dans une cellule rendent le fichier trouvable ; au passage **BUG-089** (reindex manuel ne reconstruisait pas l'index inversé)
|
||||
- [ ] **A6** Outils IA `update_xlsx_cells` / `append_xlsx_rows` / `xlsx_to_markdown` / `list_xlsx_sheets`
|
||||
- [ ] **A7** Navigation clavier + barre de formule + nom de cellule (Tab/Entrée/flèches, `Maj+Entrée`, copie de plage)
|
||||
- [ ] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse)
|
||||
- [ ] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel)
|
||||
- [ ] **A10** Types & formats de saisie (nombre/texte, booléens, dates localisées FR)
|
||||
- [x] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse) — BUG-090
|
||||
- [x] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel)
|
||||
- [x] **A10** Types & formats de saisie (nombre/texte, booléens `TRUE`/`FAUX`, dates FR `JJ/MM/AAAA` jour-first)
|
||||
- [ ] **A11** Tests frontend (`tests/frontend/xlsx-viewer.test.mjs`) + E2E (`tests/e2e/xlsx-viewer.spec.js`) au CI
|
||||
- [ ] **A12** Affichage de la valeur calculée en cache (lecture `data_only=True`, mention FR/EN)
|
||||
- [x] **A12** Valeur calculée affichée sous la formule (2ᵉ lecture `data_only=True` seulement si l'archive contient un `<v>`, info-bulle FR/EN)
|
||||
- **P2 — étendu (🟢, 2-4 j) — ⚪ à faire**
|
||||
- [ ] **A13** Tri / filtre / recherche dans la feuille + export CSV de la sélection
|
||||
- [ ] **A14** CRUD de feuilles, lignes et colonnes (renommer, insérer, supprimer, dupliquer)
|
||||
@@ -218,7 +218,7 @@
|
||||
| 🔵 Finitions | #77 Desktop : 6 tests E2E **manuels** ([protocole](./DESKTOP_E2E_CHECKLIST.md)) — signature Windows non retenue (décision 2026-09-26) | ~0,5-1 jour |
|
||||
| ⚪ P4 reporté | #73 Sync — **reporté (décision 2026-09-26)**, hors chemin critique | 6-8 jours si réactivé |
|
||||
| ⚪ P0/P1 prioritaire | #87 CI/CD (BUG-035 → BUG-040 corrigés, #86 livré) | ~3-5 jours |
|
||||
| ⚪ P0/P1/P2 backlog | #153 Visionneuse & édition XLSX — complétude (P0 ✅ A1-A4 ; A5-A12 4-6 j, A13-A17 2-4 j) | 6-11 jours restants |
|
||||
| ⚪ P0/P1/P2 backlog | #153 Visionneuse & édition XLSX — complétude (P0 ✅ A1-A4 ; P1 ✅ A5, A8-A10, A12 — reste A6-A7 ; A13-A17 2-4 j) | 3-5 jours restants |
|
||||
| **Total chemin critique** | **#77 fin + #87** | **~4-6 jours** |
|
||||
|
||||
---
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
> **Item de roadmap :** [#153 — Visionneuse & édition XLSX — complétude](../ROADMAP.md)
|
||||
> **Origine :** #152 (visionneuse XLSX, livrée en 2.27.0 — voir
|
||||
> [archive/COMPLETED_v1-v2.md](../archive/COMPLETED_v1-v2.md))
|
||||
> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), P1/P2 restants
|
||||
> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), **A8/A9 livrés le 2026-09-28** (avec BUG-090), reste A6-A7 puis A13-A17
|
||||
> **Effort estimé :** 8-13 jours au total (P0 ✅ 2-3 j · P1 4-6 j · P2 2-4 j)
|
||||
> **Règle de maintenance :** la Roadmap porte les cases à cocher (suivi), cette fiche porte
|
||||
> l'analyse, les risques et les critères d'acceptation. **Ne pas dupliquer le détail.**
|
||||
@@ -15,8 +15,9 @@
|
||||
| Couche | Fichier | Rôle |
|
||||
|---|---|---|
|
||||
| Lecture | `backend/xlsx_reader.py` | `render_sheets()` → un tableau HTML par feuille (openpyxl `read_only=True`, `data_only=False`) |
|
||||
| Endpoint lecture | `backend/routers/files_read.py:241-265` | `GET /api/file/{vault}?path=…` → `is_xlsx: true` + `xlsx_sheets: [{name, html}]` |
|
||||
| Schéma API | `backend/schemas.py:286-290` | `is_xlsx`, `xlsx_sheets` |
|
||||
| Endpoint lecture | `backend/routers/files_read.py:241-265` | `GET /api/file/{vault}?path=…` → `is_xlsx: true` + `xlsx_sheets: [{name, html, rows, cols, total_*, max_*, truncated}]` |
|
||||
| Endpoint fenêtre | `backend/routers/files_read.py` | `GET /api/file/{vault}/xlsx/sheet?path=&sheet=&offset=&limit=` (#153 A9) — une fenêtre de lignes, vraies coordonnées A1 |
|
||||
| Schéma API | `backend/schemas.py:286-290` | `is_xlsx`, `xlsx_sheets`, `XlsxSheetWindowResponse` |
|
||||
| Écriture | `backend/services/mutations.py:227-320` | `edit_xlsx_cells()` (backup, refs A1 validées, coercion `str`→`int`/`float`) |
|
||||
| Endpoint écriture | `backend/routers/files_write.py:116-148` | `PUT /api/file/{vault}/xlsx/save` (1 à 500 cellules / requête) |
|
||||
| Documentation API | `backend/openapi_docs.py:184-187` | exemple d'appel `xlsx/save` |
|
||||
@@ -105,7 +106,7 @@ restent à faire (A7).
|
||||
| R2 | Écriture non atomique (`wb.save()` en place) → classeur corrompu si crash | `mutations.edit_xlsx_cells` | **A2** — `.tmp` + `os.replace` | 🟢 livré (BUG-086) |
|
||||
| R3 | Concurrence : deux éditions (onglets, watcher + IA) → dernier écrivain gagne | `mutations.edit_xlsx_cells` | **A3** — verrou par chemin, **409** `conflict` | 🟢 livré (BUG-087) |
|
||||
| R4 | **Injection de formule** : une saisie `=cmd\|…`, `=HYPERLINK(…)` est stockée comme formule par openpyxl → DDE à l'ouverture dans Excel | `mutations._write_cell` | **A4** — forçage texte (`data_type="s"`), opt-in `allow_formula` | 🟢 livré (BUG-088) |
|
||||
| R5 | Troncature silencieuse au-delà de 500×40 | `xlsx_reader.MAX_ROWS/MAX_COLS` | A8 / A9 | ⚪ à faire |
|
||||
| R5 | Troncature silencieuse au-delà de 500×40 | `xlsx_reader.MAX_ROWS/MAX_COLS` | A8 / A9 | 🟢 bandeau + dimensions exposées (BUG-090) ; le chargement paresseux par fenêtres sert les lignes au-delà du plafond |
|
||||
|
||||
## 5. Backlog #153 — sous-tâches
|
||||
|
||||
@@ -134,13 +135,14 @@ couverture) · effort en jours-homme de développement + tests.
|
||||
nombres. Au passage : le handler `ServiceError` expose `code` + `details` et `api()` les
|
||||
propage sur l'Error. *Vérifié :* `TestXlsxFormulaGuard` (4) + test du toggle côté UI.
|
||||
|
||||
### P1 — Recherche, IA, UX (4-6 j)
|
||||
### P1 — Recherche, IA, UX (4-6 j) — 🟢 A5, A10, A12 livrés le 2026-09-28
|
||||
|
||||
- [ ] **A5 — Indexation du contenu des feuilles.** Extraire un texte (noms de feuilles +
|
||||
en-têtes + N premières lignes, plafond ~5 k caractères) pour le TF-IDF et la recherche
|
||||
sémantique, tout en gardant la lecture binaire pour l'affichage ; `content_preview`
|
||||
renseigné ; exclusion si le classeur est chiffré/corrompu. *Critère :* une cellule contenant
|
||||
un mot-clé rend le fichier trouvable ; `test_xlsx_indexing` étendu.
|
||||
- [x] **A5 — Indexation du contenu des feuilles.** `extract_indexable_text()` (noms de feuilles +
|
||||
20 premières lignes, `MAX_INDEX_CHARS = 5 000`, 20 feuilles max) alimente le TF-IDF et la
|
||||
recherche sémantique ; la lecture binaire reste inchangée pour l'affichage. Un classeur
|
||||
chiffré/corrompu s'indexe par son seul nom (jamais d'exception). Au passage : **BUG-089**,
|
||||
un reindex manuel ne reconstruisait pas l'index inversé. *Vérifié :* `TestXlsxSearchable` (4)
|
||||
+ `TestXlsxIndexing`, **contre-preuve** (neutraliser l'extraction → 3 tests échouent).
|
||||
- [ ] **A6 — Outils IA sur classeur.** `update_xlsx_cells` (enveloppe du service existant),
|
||||
`append_xlsx_rows`, `xlsx_to_markdown` (contexte LLM, plafonné), `list_xlsx_sheets` — risque
|
||||
WRITE + confirmation pour les mutations, libellés i18n dans `backend/tools/labels.py`,
|
||||
@@ -148,22 +150,42 @@ couverture) · effort en jours-homme de développement + tests.
|
||||
- [ ] **A7 — Navigation clavier & barre de formule.** `Tab`/`Maj+Tab`/`Entrée`/flèches, cellule
|
||||
active affichée (nom A1), `Maj+Entrée` pour le multiligne, copier une plage, focus visible
|
||||
et compatible mobile (≥ 44 px, `tests/e2e/mobile-editor.spec.js`).
|
||||
- [ ] **A8 — `thead` sticky + indicateur de troncature (R5).** Ligne d'en-têtes figlée au
|
||||
défilement vertical ; bandeau « feuille tronquée à 500 lignes × 40 colonnes » ; libellés
|
||||
FR/EN.
|
||||
- [ ] **A9 — Chargement paresseux par feuille (supprime le plafond).** Endpoint
|
||||
`GET /api/file/{vault}/xlsx/sheet?sheet=N&offset=&limit=` (`response_model` +
|
||||
`backend/openapi_docs.py`), rendu à la demande avec défilement virtuel, bouton « charger
|
||||
tout ».
|
||||
- [ ] **A10 — Types et formats de saisie.** Coercion symétrique à l'écriture/à l'affichage
|
||||
(nombre vs texte, booléens `TRUE`/`FALSE`, dates localisées FR — le TODO existe déjà dans
|
||||
`_coerce_xlsx_value`) ; affichage du type d'origine dans l'info-bulle de cellule.
|
||||
- [x] **A8 — `thead` sticky + indicateur de troncature (R5) — livré 2026-09-28 (BUG-090).**
|
||||
Ligne d'en-têtes figlée au défilement vertical (`thead th { top: 0 }` ; `top: auto` sur les
|
||||
numéros de ligne, sans quoi ils s'empilent en haut à gauche) ; `render_sheets()` expose
|
||||
`total_rows`/`total_cols` (dimensions déclarées), `max_rows`/`max_cols` (plafonds) et
|
||||
`truncated` — le bandeau « feuille tronquée » annonce le **plafond atteint** et non la
|
||||
taille élaguée (une feuille creuse rend 1×1 tout en couvrant 500 lignes) ; libellés
|
||||
`xlsx.truncated_*` FR/EN. *Vérifié :* `TestXlsxTruncationNotice` (4), `xlsx-viewer.test.mjs`
|
||||
(4 nouveaux), E2E sur `test_vault/sample-xlsx-large.xlsx` (520 lignes).
|
||||
- [x] **A9 — Chargement paresseux par feuille (côté API).** Endpoint
|
||||
`GET /api/file/{vault}/xlsx/sheet?sheet=&offset=&limit=` (`XlsxSheetWindowResponse`,
|
||||
exemple dans `backend/openapi_docs.py`) : une fenêtre de 1 à 1 000 lignes (plafond
|
||||
`MAX_WINDOW_ROWS`, `limit>1000` → 422), `has_more` pour paginer, valeurs calculées A12
|
||||
incluses. Les numéros de ligne et `data-cell` restent les coordonnées A1 réelles de la
|
||||
feuille (`_table(..., row_offset=offset)`) : une fenêtre est indistinguishable d'un rendu
|
||||
complet et une édition dans la fenêtre cible la bonne cellule. Erreurs : 404 feuille
|
||||
inconnue / fichier absent, 415 non-`.xlsx`. *Vérifié :* `TestXlsxSheetWindow` (11),
|
||||
**contre-preuve** (neutraliser l'offset → 3 tests échouent), E2E « l'endpoint de fenêtre
|
||||
sert les lignes au-delà du plafond ». *Reste :* défilement virtuel côté UI + bouton «
|
||||
charger tout » (le viewer garde son rendu complet ≤ 500×40, mais le bandeau A8 dit la
|
||||
vérité) ; suivi dans la Roadmap.
|
||||
- [x] **A10 — Types et formats de saisie.** `_coerce_xlsx_value()` reconnait les booléens
|
||||
(`true`/`vrai`/`oui`/`yes` et leurs négatifs) et les dates FR `JJ/MM/AAAA` (+ `HH:MM`),
|
||||
jour-first comme Excel en locale française : `01/02/2026` = 1ᵉʳ février. Une saisie
|
||||
ressemblant à une formule n'est jamais convertie (BUG-088 préservé) ; un code postal
|
||||
numérique ou une version restent ce qu'ils sont. *Vérifié :* `TestXlsxValueCoercion` (5),
|
||||
**contre-preuve** (neutraliser la coercion → 2 tests échouent).
|
||||
- [ ] **A11 — Tests frontend + E2E.** `tests/frontend/xlsx-viewer.test.mjs` (dirty, Échap,
|
||||
collage, 1 PUT par feuille, bouton désactivé) et `tests/e2e/xlsx-viewer.spec.js`
|
||||
(ouverture, onglets, édition, sauvegarde, rechargement) ; intégration au CI.
|
||||
- [ ] **A12 — Valeurs calculées.** Afficher la valeur en cache (2ᵉ ligne discrète) quand elle
|
||||
existe, via une lecture `data_only=True` de la même page d'onglets ; mention FR/EN
|
||||
« valeur recalculée par Excel ».
|
||||
- [x] **A12 — Valeurs calculées.** La valeur en cache s'affiche sous la formule dans un
|
||||
`<span class="xlsx-cached">`. La 2ᵉ lecture `data_only=True` n'a lieu que si l'archive
|
||||
contient réellement un `<f>…</f><v>…</v>` (sonde déjà présente pour A1) : le cas courant
|
||||
reste à un seul chargement, et toute erreur retombe sur l'affichage formules seul.
|
||||
L'info-bulle est traduite côté client (`xlsx.cached_value_title` FR/EN) — aucun texte
|
||||
d'interface n'est émis par le backend. *Vérifié :* `TestXlsxCachedValues` (3),
|
||||
**contre-preuve** (neutraliser la 2ᵉ lecture → 2 tests échouent).
|
||||
|
||||
### P2 — Étendu (2-4 j)
|
||||
|
||||
@@ -199,3 +221,5 @@ couverture) · effort en jours-homme de développement + tests.
|
||||
| 2026-09-27 | Audit complet → création de #153 : limites, risques R1-R5, backlog A1-A17 |
|
||||
| 2026-09-27 | Périmètre de perte **remesuré** sur openpyxl 3.1.5 : graphiques / images / TCD sont préservés, seules les valeurs en cache et quelques parties exotiques sont perdues |
|
||||
| 2026-09-27 | **P0 livré** (BUG-085 → BUG-088) : `xlsx_lossy_features` + 409 `xlsx_lossy_content`, écriture atomique, verrou par fichier, formules stockées en texte par défaut |
|
||||
| 2026-09-28 | **A5 + A10 + A12 livrés** : le contenu des cellules est indexé (recherche), la saisie est typée (booléens, dates FR), la valeur calculée s'affiche sous la formule. **BUG-089** corrigé au passage (reindex manuel ≠ reconstruction de l'index inversé ; `backend/search.py` lisait l'index par valeur) |
|
||||
| 2026-09-28 | **A8 + A9 livrés** (BUG-090) : la troncature d'une feuille est annoncée (bandeau + dimensions dans la réponse de lecture), les en-têtes restent visibles au défilement, et `GET …/xlsx/sheet` sert une fenêtre de lignes avec les vraies coordonnées A1 — les lignes au-delà du plafond redeviennent accessibles aux clients API. Défilement virtuel côté UI à suivre |
|
||||
|
||||
+36
-1
@@ -1005,6 +1005,35 @@ export function renderVideoViewer(area, data) {
|
||||
// confirmation before retrying with `force: true`. A value starting with "=" or
|
||||
// "@" is stored as text unless the user turns the formula toggle on, so a typed
|
||||
// `=cmd|…` cannot execute when the file is later opened in Excel.
|
||||
// #153 A8 — a sheet bigger than the render caps used to be silently cut: the
|
||||
// user saw a short table and no way to tell the rest of the workbook still
|
||||
// existed. The backend now reports the real dimensions of every sheet, so the
|
||||
// note states exactly what is hidden (and that those cells are not editable
|
||||
// here — the workbook itself is untouched). A payload without those fields
|
||||
// (older cache) simply shows no note.
|
||||
function truncationNote(sheet) {
|
||||
// The cap is the real "shown" figure, not `rows`/`cols`: those are post-trim
|
||||
// (a sparse sheet renders 1x1) while the note must say how far the view
|
||||
// reaches.
|
||||
const cap = { rows: Number(sheet.max_rows) || 0, cols: Number(sheet.max_cols) || 0 };
|
||||
const reasons = [];
|
||||
if (Number(sheet.total_rows) > cap.rows) {
|
||||
reasons.push(t("xlsx.truncated_rows", { shown: cap.rows, total: Number(sheet.total_rows) }));
|
||||
}
|
||||
if (Number(sheet.total_cols) > cap.cols) {
|
||||
reasons.push(t("xlsx.truncated_cols", { shown: cap.cols, total: Number(sheet.total_cols) }));
|
||||
}
|
||||
if (!reasons.length) return "";
|
||||
return `<div class="xlsx-truncated" role="note">
|
||||
<i data-lucide="scissors" class="xlsx-truncated-icon"></i>
|
||||
<div class="xlsx-warning-body">
|
||||
<strong>${escapeHtml(t("xlsx.truncated_title"))}</strong>
|
||||
<span>${escapeHtml(reasons.join(" "))}</span>
|
||||
<span class="xlsx-warning-hint">${escapeHtml(t("xlsx.truncated_hint"))}</span>
|
||||
</div>
|
||||
</div>`;
|
||||
}
|
||||
|
||||
export function renderXlsxViewer(area, data) {
|
||||
const sheets = data.xlsx_sheets || [];
|
||||
const lossy = data.xlsx_lossy_features || [];
|
||||
@@ -1019,7 +1048,7 @@ export function renderXlsxViewer(area, data) {
|
||||
).join("")}</div>`
|
||||
: "";
|
||||
const panels = sheets.map((s, i) =>
|
||||
`<div class="xlsx-panel" data-sheet="${i}"${i === 0 ? "" : ' style="display:none"'}>${s.html}</div>`
|
||||
`<div class="xlsx-panel" data-sheet="${i}"${i === 0 ? "" : ' style="display:none"'}>${truncationNote(s)}${s.html}</div>`
|
||||
).join("");
|
||||
const lossWarning = lossy.length
|
||||
? `<div class="xlsx-warning" role="note">
|
||||
@@ -1056,6 +1085,12 @@ export function renderXlsxViewer(area, data) {
|
||||
const dirtyCount = () => area.querySelectorAll("td.xlsx-dirty").length;
|
||||
const refreshSaveState = () => { saveBtn.disabled = dirtyCount() === 0; };
|
||||
|
||||
// #153 A12 — the backend marks the last value Excel computed; the wording is
|
||||
// translated here so the tooltip follows the UI language.
|
||||
area.querySelectorAll(".xlsx-cached[data-cached-value]").forEach((el) => {
|
||||
el.title = t("xlsx.cached_value_title");
|
||||
});
|
||||
|
||||
// Editable cells: Enter blurs, Escape reverts, paste stays single-line.
|
||||
area.querySelectorAll(".xlsx-table td").forEach((td) => {
|
||||
td.contentEditable = "true";
|
||||
|
||||
@@ -1828,6 +1828,11 @@
|
||||
"xlsx.lossy_confirm": "Save anyway? The following will be lost: {features}",
|
||||
"xlsx.lossy_cancelled": "Save cancelled",
|
||||
"xlsx.formula_toggle_title": "Treat “=” and “@” as formulas (off by default)",
|
||||
"xlsx.cached_value_title": "Last value calculated by Excel",
|
||||
"xlsx.truncated_title": "Truncated sheet",
|
||||
"xlsx.truncated_rows": "{shown} of {total} rows displayed.",
|
||||
"xlsx.truncated_cols": "{shown} of {total} columns displayed.",
|
||||
"xlsx.truncated_hint": "Cells outside the displayed area cannot be edited here; the workbook is unchanged.",
|
||||
"xlsx.feature_cached_values": "cached values",
|
||||
"xlsx.feature_slicers": "slicers and timelines",
|
||||
"xlsx.feature_form_controls": "form controls",
|
||||
|
||||
@@ -1828,6 +1828,11 @@
|
||||
"xlsx.lossy_confirm": "Enregistrer quand même ? Les éléments suivants seront perdus : {features}",
|
||||
"xlsx.lossy_cancelled": "Sauvegarde annulée",
|
||||
"xlsx.formula_toggle_title": "Interpréter « = » et « @ » comme des formules (désactivé par défaut)",
|
||||
"xlsx.cached_value_title": "Dernière valeur calculée par Excel",
|
||||
"xlsx.truncated_title": "Feuille tronquée",
|
||||
"xlsx.truncated_rows": "{shown} lignes affichées sur {total}.",
|
||||
"xlsx.truncated_cols": "{shown} colonnes affichées sur {total}.",
|
||||
"xlsx.truncated_hint": "Les cellules hors de l'affichage ne sont pas éditables ici ; le classeur n'est pas modifié.",
|
||||
"xlsx.feature_cached_values": "valeurs calculées",
|
||||
"xlsx.feature_slicers": "segments et chronologies",
|
||||
"xlsx.feature_form_controls": "contrôles de formulaire",
|
||||
|
||||
+54
-2
@@ -10967,12 +10967,25 @@ body.desktop-mode .editor-container {
|
||||
border-right: 1px solid var(--border-light, var(--border));
|
||||
position: sticky;
|
||||
left: 0;
|
||||
z-index: 1;
|
||||
/* #153 A8 — `top: auto` is load-bearing: `.csv-table th` pins EVERY `th`
|
||||
at `top: 0`, so a row number left sticky on both axes piles up in the
|
||||
top-left corner instead of tracking its own row. */
|
||||
top: auto;
|
||||
z-index: 2;
|
||||
}
|
||||
.xlsx-table th.xlsx-corner {
|
||||
left: 0;
|
||||
top: 0;
|
||||
z-index: 2;
|
||||
z-index: 4;
|
||||
}
|
||||
/* #153 A8 — the column headers stay visible while the sheet scrolls down.
|
||||
Declared explicitly (and not inherited from `.csv-table th`) so the stacking
|
||||
order is intentional: thead (3) < row numbers (2) < corner (4). */
|
||||
.xlsx-table thead th {
|
||||
position: sticky;
|
||||
top: 0;
|
||||
z-index: 3;
|
||||
background: var(--surface);
|
||||
}
|
||||
.xlsx-table td[contenteditable] {
|
||||
cursor: text;
|
||||
@@ -10987,6 +11000,20 @@ body.desktop-mode .editor-container {
|
||||
background: rgba(255, 196, 0, 0.18);
|
||||
}
|
||||
|
||||
/* #153 A12 — last result Excel computed, shown under a formula cell.
|
||||
Discreet by design: the formula is what the user edits, the cached value is
|
||||
context (stale until Excel recalculates). */
|
||||
.xlsx-cached {
|
||||
display: block;
|
||||
margin-top: 2px;
|
||||
padding-left: 6px;
|
||||
border-left: 2px solid var(--border, #d0d7de);
|
||||
color: var(--text-muted);
|
||||
font-size: 0.85em;
|
||||
font-variant-numeric: tabular-nums;
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
/* #153 A1/A4 — lossy-save warning + formula toggle */
|
||||
.xlsx-warning {
|
||||
display: flex;
|
||||
@@ -11037,6 +11064,31 @@ body.desktop-mode .editor-container {
|
||||
color: var(--text-secondary);
|
||||
opacity: 0.85;
|
||||
}
|
||||
/* #153 A8 — "feuille tronquée" notice. Deliberately NOT the `.xlsx-warning`
|
||||
look: that one is a data-loss alert, this one only says part of the sheet is
|
||||
out of view. */
|
||||
.xlsx-truncated {
|
||||
display: flex;
|
||||
align-items: flex-start;
|
||||
gap: 8px;
|
||||
padding: 8px 10px;
|
||||
margin-bottom: 8px;
|
||||
border: 1px solid var(--border);
|
||||
border-left: 3px solid var(--accent, #4a90d9);
|
||||
border-radius: 4px;
|
||||
background: var(--bg-secondary);
|
||||
color: var(--text-secondary);
|
||||
font-size: 0.82rem;
|
||||
line-height: 1.45;
|
||||
}
|
||||
.xlsx-truncated-icon {
|
||||
width: 16px;
|
||||
height: 16px;
|
||||
flex: 0 0 auto;
|
||||
margin-top: 1px;
|
||||
color: var(--accent, #4a90d9);
|
||||
}
|
||||
|
||||
.xlsx-formula-toggle {
|
||||
font-family: 'JetBrains Mono', 'Fira Code', 'Consolas', monospace;
|
||||
font-weight: 600;
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "obsigate",
|
||||
"version": "2.29.0",
|
||||
"version": "2.31.0",
|
||||
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
|
||||
"main": "patch.js",
|
||||
"directories": {
|
||||
|
||||
Binary file not shown.
@@ -11,7 +11,14 @@
|
||||
* - the warning banner lists both features ;
|
||||
* - saving a cell on that workbook asks for confirmation (native dialog) and
|
||||
* then succeeds (the client retries with `force: true`) ;
|
||||
* - the f(x) toggle is off by default, so "=B1*3" is stored as text.
|
||||
* - the f(x) toggle is off by default, so "=B1*3" is stored as text ;
|
||||
* - the value Excel last computed is shown under the formula (#153 A12).
|
||||
*
|
||||
* Second describe block — `test_vault/sample-xlsx-large.xlsx` (520 rows) :
|
||||
* - a sheet over the render caps SAYS it instead of looking complete (#153 A8) ;
|
||||
* - the column headers stay pinned while the sheet scrolls (#153 A8) ;
|
||||
* - `GET …/xlsx/sheet?offset=500` serves the rows the caps used to hide,
|
||||
* with the real A1 coordinates (#153 A9).
|
||||
*
|
||||
* The fixture is restored byte-for-byte in `afterAll` so a local run never
|
||||
* dirties the working copy.
|
||||
@@ -26,6 +33,10 @@ import path from 'node:path';
|
||||
const BASE = process.env.BASE_URL || 'http://localhost:2029';
|
||||
const VAULT = 'TestVault';
|
||||
const FIXTURE = 'sample-xlsx-lossy.xlsx';
|
||||
// #153 A8 — 520 rows x 3 columns: the sheet exceeds the 500-row render cap, so
|
||||
// the viewer must SAY so. Generated once with openpyxl (header + 519 lines) and
|
||||
// committed next to the other fixture; nothing in the suite writes to it.
|
||||
const LARGE = 'sample-xlsx-large.xlsx';
|
||||
// Playwright runs from the repository root (run-e2e-local.* / CI both do).
|
||||
const FIXTURE_PATH = path.resolve(process.cwd(), 'test_vault', FIXTURE);
|
||||
|
||||
@@ -44,7 +55,11 @@ async function login(page) {
|
||||
}
|
||||
|
||||
async function openFixture(page) {
|
||||
const treeItem = page.locator(`.tree-item[data-vault="${VAULT}"][data-path="${FIXTURE}"]`);
|
||||
return openXlsx(page, FIXTURE);
|
||||
}
|
||||
|
||||
async function openXlsx(page, file) {
|
||||
const treeItem = page.locator(`.tree-item[data-vault="${VAULT}"][data-path="${file}"]`);
|
||||
if (!(await treeItem.count())) {
|
||||
await page.locator(`.tree-item.vault-item[data-vault="${VAULT}"]`).first().click();
|
||||
await treeItem.waitFor({ state: 'attached', timeout: 8000 });
|
||||
@@ -53,7 +68,7 @@ async function openFixture(page) {
|
||||
await expect(page.locator('#content-area .xlsx-table')).toBeVisible({ timeout: 15000 });
|
||||
}
|
||||
|
||||
test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
|
||||
test.describe('Excel viewer — garde-fous d\'écriture et valeurs calculées (#153)', () => {
|
||||
test.beforeAll(() => {
|
||||
if (existsSync(FIXTURE_PATH)) originalBytes = readFileSync(FIXTURE_PATH);
|
||||
});
|
||||
@@ -62,6 +77,28 @@ test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
|
||||
if (originalBytes) writeFileSync(FIXTURE_PATH, originalBytes);
|
||||
});
|
||||
|
||||
// Read-only assertions come FIRST, before the mutating tests: saving through
|
||||
// the viewer rewrites the workbook and an openpyxl round-trip drops the cached
|
||||
// formula results (BUG-085), so the shadow line only exists on a pristine
|
||||
// fixture.
|
||||
test('affiche la valeur calculée en cache sous la formule (#153 A12)', async ({ page }) => {
|
||||
await login(page);
|
||||
await openFixture(page);
|
||||
|
||||
// B2 is "=B1*2" and the package keeps its last result (<v>200</v>).
|
||||
const formulaCell = page.locator('#content-area td[data-cell="B2"]');
|
||||
await expect(formulaCell).toContainText('=B1*2');
|
||||
|
||||
const cached = formulaCell.locator('.xlsx-cached');
|
||||
await expect(cached).toHaveCount(1);
|
||||
await expect(cached).toHaveText('200');
|
||||
// The tooltip is translated client-side, never hardcoded by the backend.
|
||||
await expect(cached).toHaveAttribute('title', /Excel/);
|
||||
|
||||
// A plain value cell must not be duplicated with a shadow line.
|
||||
await expect(page.locator('#content-area td[data-cell="B1"] .xlsx-cached')).toHaveCount(0);
|
||||
});
|
||||
|
||||
test('affiche la bannière listant les éléments non préservés', async ({ page }) => {
|
||||
await login(page);
|
||||
await openFixture(page);
|
||||
@@ -118,3 +155,59 @@ test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
|
||||
await expect(page.locator('.toast-success')).toBeVisible({ timeout: 10000 });
|
||||
});
|
||||
});
|
||||
|
||||
// ── A8 — troncature annoncée + en-têtes figés ───────────────────────────────
|
||||
|
||||
test.describe('Excel viewer — troncature et navigation (#153 A8/A9)', () => {
|
||||
test('annonce la feuille tronquée au lieu de la couper en silence', async ({ page }) => {
|
||||
await login(page);
|
||||
await openXlsx(page, LARGE);
|
||||
|
||||
const note = page.locator('#content-area .xlsx-truncated');
|
||||
await expect(note).toBeVisible();
|
||||
// Libellé traduit (jamais de texte UI backend, jamais de couleur en dur).
|
||||
await expect(note).toContainText('Feuille tronquée');
|
||||
await expect(note).toContainText('500 lignes affichées sur 520');
|
||||
|
||||
// La dernière ligne rendue est la 500e ; les suivantes ne sont pas là.
|
||||
await expect(page.locator('#content-area td[data-cell="A500"]')).toHaveCount(1);
|
||||
await expect(page.locator('#content-area td[data-cell="A501"]')).toHaveCount(0);
|
||||
});
|
||||
|
||||
test('garde les en-têtes de colonnes visibles au défilement', async ({ page }) => {
|
||||
await login(page);
|
||||
await openXlsx(page, LARGE);
|
||||
|
||||
const header = page.locator('#content-area .xlsx-table thead th').nth(1);
|
||||
const before = await header.boundingBox();
|
||||
|
||||
await page.locator('#content-area .csv-table-wrapper').evaluate((el) => { el.scrollTop = 800; });
|
||||
await expect.poll(async () => (await header.boundingBox()).y, { timeout: 5000 })
|
||||
.toBeLessThanOrEqual(before.y + 1);
|
||||
|
||||
// Les numéros de ligne ne se superposent pas en haut à gauche (le `top: auto`
|
||||
// de A8) et la première ligne de données reste lisible sous l'en-tête.
|
||||
const first = await page.locator('#content-area th.xlsx-rownum').first().boundingBox();
|
||||
const second = await page.locator('#content-area th.xlsx-rownum').nth(1).boundingBox();
|
||||
expect(second.y - first.y).toBeGreaterThan(4);
|
||||
});
|
||||
|
||||
test('l\'endpoint de fenêtre sert les lignes au-delà du plafond (#153 A9)', async ({ page }) => {
|
||||
await login(page);
|
||||
await openXlsx(page, LARGE);
|
||||
|
||||
const res = await page.request.get(
|
||||
`${BASE}/api/file/${VAULT}/xlsx/sheet?path=${encodeURIComponent(LARGE)}&sheet=Journal&offset=500&limit=50`
|
||||
);
|
||||
expect(res.status()).toBe(200);
|
||||
const win = await res.json();
|
||||
expect(win.total_rows).toBe(520);
|
||||
expect(win.offset).toBe(500);
|
||||
expect(win.has_more).toBe(false);
|
||||
// Les coordonnées A1 sont celles de la feuille, pas celles de la fenêtre :
|
||||
// la ligne 520 est servie comme A520, pas comme A20.
|
||||
expect(win.html).toContain('data-cell="A520"');
|
||||
expect(win.html).toContain('Operation 519');
|
||||
expect(win.html).not.toContain('data-cell="A1"');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -7,6 +7,7 @@
|
||||
* workbook gets 409 `xlsx_lossy_content`, asks for confirmation and
|
||||
* retries with `force: true` (or gives up when refused);
|
||||
* - A4 : the f(x) toggle flips `allow_formula` in the save payload.
|
||||
* - A8 : a sheet bigger than the render caps shows the truncation notice.
|
||||
*
|
||||
* Usage: node tests/frontend/xlsx-viewer.test.mjs
|
||||
*/
|
||||
@@ -116,14 +117,14 @@ const sheetHtml = (value) =>
|
||||
`<tr><th class="xlsx-rownum">1</th><td data-cell="A1">${value}</td></tr>` +
|
||||
"</tbody></table></div>";
|
||||
|
||||
function mount({ lossy = [] } = {}) {
|
||||
function mount({ lossy = [], sheet = {} } = {}) {
|
||||
const area = document.getElementById("content-area");
|
||||
area.innerHTML = "";
|
||||
renderXlsxViewer(area, {
|
||||
vault: "V",
|
||||
path: "data.xlsx",
|
||||
is_xlsx: true,
|
||||
xlsx_sheets: [{ name: "Feuille1", html: sheetHtml("100") }],
|
||||
xlsx_sheets: [{ name: "Feuille1", html: sheetHtml("100"), ...sheet }],
|
||||
xlsx_lossy_features: lossy,
|
||||
});
|
||||
return area;
|
||||
@@ -268,6 +269,49 @@ await test("a non-409 failure is not retried", async () => {
|
||||
assert.equal(area.querySelectorAll("td.xlsx-dirty").length, 1);
|
||||
});
|
||||
|
||||
// ── A8 — truncation notice ───────────────────────────────────────────────────
|
||||
|
||||
await test("no notice when the sheet fits within the render caps", () => {
|
||||
const area = mount({
|
||||
sheet: { rows: 500, cols: 40, total_rows: 500, total_cols: 40, max_rows: 500, max_cols: 40, truncated: false },
|
||||
});
|
||||
assert.equal(area.querySelector(".xlsx-truncated"), null);
|
||||
});
|
||||
|
||||
await test("notice states the cap and the real size of a truncated sheet", () => {
|
||||
const area = mount({
|
||||
sheet: { rows: 500, cols: 12, total_rows: 1200, total_cols: 12, max_rows: 500, max_cols: 40, truncated: true },
|
||||
});
|
||||
const note = area.querySelector(".xlsx-truncated");
|
||||
assert.ok(note, "notice absent");
|
||||
const txt = note.textContent;
|
||||
assert.ok(txt.includes(FR["xlsx.truncated_title"]), txt);
|
||||
// {shown} is the CAP (500), not the post-trim row count: a sparse sheet
|
||||
// renders 1 row but the view still reaches 500 of them.
|
||||
const expected = FR["xlsx.truncated_rows"].replace("{shown}", "500").replace("{total}", "1200");
|
||||
assert.ok(txt.includes(expected), `${txt} !includes ${expected}`);
|
||||
// Nothing to say about the columns here (12 < 40).
|
||||
assert.ok(!txt.includes(FR["xlsx.truncated_cols"]), txt);
|
||||
});
|
||||
|
||||
await test("notice mentions both axes when rows AND columns overflow", () => {
|
||||
const area = mount({
|
||||
sheet: { rows: 1, cols: 40, total_rows: 501, total_cols: 45, max_rows: 500, max_cols: 40, truncated: true },
|
||||
});
|
||||
const txt = area.querySelector(".xlsx-truncated").textContent;
|
||||
assert.ok(
|
||||
txt.includes(FR["xlsx.truncated_cols"].replace("{shown}", "40").replace("{total}", "45")),
|
||||
txt
|
||||
);
|
||||
});
|
||||
|
||||
await test("a payload without the dimensions shows no notice", () => {
|
||||
// Backward compatibility: an older cached response must not produce "NaN".
|
||||
const area = mount({ sheet: { name: "Feuille1" } });
|
||||
assert.equal(area.querySelector(".xlsx-truncated"), null);
|
||||
assert.ok(!area.textContent.includes("NaN"));
|
||||
});
|
||||
|
||||
// ── Report ──────────────────────────────────────────────────────────────────
|
||||
console.log(`\n${passCount}/${testCount} tests passed\n`);
|
||||
process.exit(passCount === testCount ? 0 : 1);
|
||||
|
||||
+361
-2
@@ -106,16 +106,208 @@ class TestXlsxIndexing:
|
||||
|
||||
assert ".xlsx" in SUPPORTED_EXTENSIONS
|
||||
|
||||
def test_xlsx_indexed_metadata_only(self, test_vault_dir, xlsx_file):
|
||||
def test_xlsx_indexes_sheet_names_and_headers(self, test_vault_dir, xlsx_file):
|
||||
"""#153 A5 — a workbook is searchable by its cell values."""
|
||||
from backend.indexer import _index_single_file_sync
|
||||
|
||||
info = _index_single_file_sync(VAULT, test_vault_dir, xlsx_file)
|
||||
assert info is not None
|
||||
assert info["extension"] == ".xlsx"
|
||||
assert info["content"] == "" # binary: never read into TF-IDF
|
||||
# Sheet names + header rows reach TF-IDF (was metadata-only before A5).
|
||||
assert "Budget" in info["content"]
|
||||
assert "Poste" in info["content"]
|
||||
assert info["content_preview"]
|
||||
assert info["title"] # filename-derived title
|
||||
|
||||
|
||||
class TestXlsxCachedValues:
|
||||
"""#153 A12 — show what Excel last computed next to each formula."""
|
||||
|
||||
def test_cached_result_is_shown_beside_the_formula(self, client, lossy_xlsx):
|
||||
from backend.xlsx_reader import render_sheets
|
||||
|
||||
html = render_sheets(Path(lossy_xlsx))[0]["html"]
|
||||
# A2 is "=A1*3" with <v>9</v> in the fixture.
|
||||
assert 'data-cell="A2"' in html
|
||||
assert "=A1*3" in html
|
||||
assert "xlsx-cached" in html
|
||||
assert ">9<" in html # the cached result Excel computed
|
||||
|
||||
def test_no_shadow_when_no_formula_carries_a_result(self, client, xlsx_file):
|
||||
from backend.xlsx_reader import render_sheets
|
||||
|
||||
html = render_sheets(Path(xlsx_file))[0]["html"]
|
||||
assert "xlsx-cached" not in html # budget.xlsx has no <v> at all
|
||||
|
||||
def test_plain_cells_are_never_duplicated(self, client, lossy_xlsx):
|
||||
from backend.xlsx_reader import render_sheets
|
||||
|
||||
html = render_sheets(Path(lossy_xlsx))[0]["html"]
|
||||
# A1 is the literal 3: the two reads agree, so only one value shows.
|
||||
assert 'data-cell="A1"' in html
|
||||
assert html.count("xlsx-cached") == 1 # only the formula cell
|
||||
|
||||
|
||||
class TestXlsxValueCoercion:
|
||||
"""#153 A10 — a typed value comes back with the type Excel would infer."""
|
||||
|
||||
def _write(self, client, xlsx_file, ref, value):
|
||||
return client.put(
|
||||
f"/api/file/{VAULT}/xlsx/save",
|
||||
params={"path": "budget.xlsx"},
|
||||
json={"sheet": "Budget", "cells": {ref: value}, "force": True},
|
||||
)
|
||||
|
||||
def test_number_and_bool_are_stored_as_typed(self, client, xlsx_file):
|
||||
resp = self._write(
|
||||
client, xlsx_file, "D1", "42"
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
resp = self._write(client, xlsx_file, "D2", "VRAI")
|
||||
assert resp.status_code == 200
|
||||
resp = self._write(client, xlsx_file, "D3", "12/03/2026")
|
||||
assert resp.status_code == 200
|
||||
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(xlsx_file)
|
||||
ws = wb["Budget"]
|
||||
assert ws["D1"].value == 42 and isinstance(ws["D1"].value, int)
|
||||
assert ws["D2"].value is True
|
||||
assert ws["D3"].value.year == 2026 and ws["D3"].value.month == 3
|
||||
assert ws["D3"].value.day == 12 # FR day-first, not 3 December
|
||||
wb.close()
|
||||
|
||||
def test_day_first_date_is_not_read_as_us(self, client, xlsx_file):
|
||||
"""'01/02/2026' is 1 February in French, not 2 January."""
|
||||
assert self._write(client, xlsx_file, "E1", "01/02/2026").status_code == 200
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(xlsx_file)
|
||||
d = wb["Budget"]["E1"].value
|
||||
wb.close()
|
||||
assert (d.month, d.day) == (2, 1)
|
||||
|
||||
def test_ambiguous_text_is_left_alone(self, client, xlsx_file):
|
||||
"""A version string or a partial date stays text, never a date."""
|
||||
assert self._write(client, xlsx_file, "F1", "3.14.2").status_code == 200
|
||||
assert self._write(client, xlsx_file, "F2", "Ref 12/34").status_code == 200
|
||||
assert self._write(client, xlsx_file, "F3", "12/2026").status_code == 200
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(xlsx_file)
|
||||
ws = wb["Budget"]
|
||||
assert ws["F1"].value == "3.14.2"
|
||||
assert ws["F2"].value == "Ref 12/34"
|
||||
assert ws["F3"].value == "12/2026"
|
||||
wb.close()
|
||||
|
||||
def test_plain_integer_text_becomes_a_number(self, client, xlsx_file):
|
||||
"""Typing a bare number yields a number, as it did before #153 A10."""
|
||||
assert self._write(client, xlsx_file, "H1", "75001").status_code == 200
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(xlsx_file)
|
||||
assert wb["Budget"]["H1"].value == 75001
|
||||
wb.close()
|
||||
|
||||
def test_formula_looking_date_stays_text(self, client, xlsx_file):
|
||||
"""A formula is never mistaken for a date (BUG-088 must not regress)."""
|
||||
assert self._write(client, xlsx_file, "G1", "=12/03/2026").status_code == 200
|
||||
from openpyxl import load_workbook
|
||||
|
||||
wb = load_workbook(xlsx_file)
|
||||
c = wb["Budget"]["G1"]
|
||||
wb.close()
|
||||
assert c.data_type == "s"
|
||||
assert c.value == "=12/03/2026"
|
||||
|
||||
|
||||
class TestXlsxSearchable:
|
||||
"""#153 A5 — a keyword living in a CELL must make the file findable."""
|
||||
|
||||
def test_search_finds_a_word_stored_in_a_cell(self, test_vault_dir):
|
||||
"""A word typed in a CELL must make the workbook findable.
|
||||
|
||||
Deliberately hermetic: it drives the indexer and the inverted index
|
||||
directly instead of going through the HTTP reload, because both are
|
||||
process-wide singletons that other test modules mutate (some reload the
|
||||
module outright), which would make this test order-dependent.
|
||||
"""
|
||||
from openpyxl import Workbook
|
||||
|
||||
import backend.indexer as indexer
|
||||
import backend.search as search_mod
|
||||
|
||||
path = Path(test_vault_dir) / "fournisseurs.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "Contacts"
|
||||
ws["A1"] = "Fournisseur"
|
||||
ws["A2"] = "Menuiserie Beaulieu"
|
||||
wb.save(path)
|
||||
|
||||
# Index exactly this vault, from scratch.
|
||||
indexer.index[VAULT] = {
|
||||
"files": [indexer._index_single_file_sync(VAULT, test_vault_dir, str(path))],
|
||||
"tags": {},
|
||||
"paths": [],
|
||||
"config": {},
|
||||
}
|
||||
entries = [f for f in indexer.index[VAULT]["files"] if f]
|
||||
assert entries, "le tableur n'a pas ete indexe"
|
||||
assert "Menuiserie" in entries[0]["content"], (
|
||||
f"contenu indexe : {entries[0]['content'][:80]!r}"
|
||||
)
|
||||
|
||||
search_mod.init_inverted_index()
|
||||
hits = [r["path"] for r in search_mod.search("Menuiserie", "all")]
|
||||
assert "fournisseurs.xlsx" in hits
|
||||
|
||||
def test_indexable_text_is_capped(self, tmp_path):
|
||||
"""A data dump must not flood the index."""
|
||||
from openpyxl import Workbook
|
||||
|
||||
from backend.xlsx_reader import MAX_INDEX_CHARS, extract_indexable_text
|
||||
|
||||
path = tmp_path / "huge.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "Big"
|
||||
for r in range(1, 400):
|
||||
ws.cell(row=r, column=1, value=f"ligne {r} " + "x" * 60)
|
||||
wb.save(path)
|
||||
|
||||
text = extract_indexable_text(path)
|
||||
assert len(text) <= MAX_INDEX_CHARS
|
||||
assert "Big" in text # the sheet name survives
|
||||
|
||||
def test_corrupt_workbook_indexes_as_empty_not_a_crash(self, tmp_path):
|
||||
from backend.xlsx_reader import extract_indexable_text
|
||||
|
||||
path = tmp_path / "broken.xlsx"
|
||||
path.write_bytes(b"not a zip at all")
|
||||
assert extract_indexable_text(path) == ""
|
||||
|
||||
def test_blank_rows_are_skipped(self, tmp_path):
|
||||
from openpyxl import Workbook
|
||||
|
||||
from backend.xlsx_reader import extract_indexable_text
|
||||
|
||||
path = tmp_path / "sparse.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "S"
|
||||
ws["A1"] = "Alpha"
|
||||
ws["A50"] = "Omega" # beyond the indexed prefix -> ignored on purpose
|
||||
wb.save(path)
|
||||
|
||||
text = extract_indexable_text(path)
|
||||
assert "Alpha" in text
|
||||
assert "Omega" not in text
|
||||
assert "\t\t" not in text # no run of empty columns
|
||||
|
||||
|
||||
# ── Edit ──────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -327,6 +519,173 @@ class TestXlsxWriteLock:
|
||||
assert openpyxl.load_workbook(xlsx_file)["Budget"]["A1"].value == "premier"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def wide_xlsx(test_vault_dir: str) -> str:
|
||||
"""Workbook whose sheet exceeds BOTH render caps (501 rows x 45 cols).
|
||||
|
||||
Sparse on purpose: a cell in A501 and one in AS1 are enough for openpyxl
|
||||
to declare those dimensions, without writing 20 000 cells to disk.
|
||||
"""
|
||||
from openpyxl import Workbook
|
||||
|
||||
path = Path(test_vault_dir) / "grand.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "Data"
|
||||
ws["A1"] = "tête"
|
||||
ws["A501"] = "dernière ligne"
|
||||
ws["AS1"] = "colonne 45"
|
||||
wb.save(path)
|
||||
return str(path)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def edge_xlsx(test_vault_dir: str) -> str:
|
||||
"""Sheet exactly on the caps (500 rows x 40 cols) — must NOT be truncated."""
|
||||
from openpyxl import Workbook
|
||||
|
||||
path = Path(test_vault_dir) / "limite.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "Data"
|
||||
ws["A1"] = "bord"
|
||||
ws["A500"] = "ligne 500"
|
||||
ws["AN1"] = "colonne 40"
|
||||
wb.save(path)
|
||||
return str(path)
|
||||
|
||||
|
||||
# ── #153 A8 — silent truncation made visible ─────────────────────────────
|
||||
|
||||
|
||||
class TestXlsxTruncationNotice:
|
||||
"""A sheet bigger than the caps must SAY so instead of looking complete."""
|
||||
|
||||
def test_render_reports_the_real_dimensions(self, client, wide_xlsx):
|
||||
resp = client.get(f"/api/file/{VAULT}", params={"path": "grand.xlsx"})
|
||||
sheet = resp.json()["xlsx_sheets"][0]
|
||||
assert (sheet["total_rows"], sheet["total_cols"]) == (501, 45)
|
||||
assert sheet["truncated"] is True
|
||||
# The caps are the coverage the banner announces — `rows`/`cols` are
|
||||
# post-trim and would understate it on a sparse sheet.
|
||||
assert (sheet["max_rows"], sheet["max_cols"]) == (500, 40)
|
||||
assert (sheet["rows"], sheet["cols"]) == (1, 1) # only 3 filled cells
|
||||
|
||||
def test_a_sheet_on_the_caps_is_not_flagged(self, client, edge_xlsx):
|
||||
"""Boundary: 500x40 is exactly what the renderer supports."""
|
||||
resp = client.get(f"/api/file/{VAULT}", params={"path": "limite.xlsx"})
|
||||
sheet = resp.json()["xlsx_sheets"][0]
|
||||
assert sheet["truncated"] is False
|
||||
assert (sheet["total_rows"], sheet["total_cols"]) == (500, 40)
|
||||
|
||||
def test_a_normal_sheet_is_not_flagged(self, client, xlsx_file):
|
||||
resp = client.get(f"/api/file/{VAULT}", params={"path": "budget.xlsx"})
|
||||
assert all(not s["truncated"] for s in resp.json()["xlsx_sheets"])
|
||||
|
||||
def test_blank_tail_is_not_reported_as_truncation(self, client, test_vault_dir):
|
||||
"""A sheet with empty rows below its data fits in the caps."""
|
||||
from openpyxl import Workbook
|
||||
|
||||
path = Path(test_vault_dir) / "blanc.xlsx"
|
||||
wb = Workbook()
|
||||
ws = wb.active
|
||||
ws.title = "Data"
|
||||
ws["A1"] = "seule ligne"
|
||||
ws["A300"] = None # formatted-but-empty row inside the caps
|
||||
wb.save(path)
|
||||
resp = client.get(f"/api/file/{VAULT}", params={"path": "blanc.xlsx"})
|
||||
sheet = resp.json()["xlsx_sheets"][0]
|
||||
assert sheet["truncated"] is False
|
||||
assert sheet["rows"] == 1 # trailing blanks dropped by _trim
|
||||
|
||||
|
||||
# ── #153 A9 — lazy per-sheet loading ─────────────────────────────────────
|
||||
|
||||
|
||||
class TestXlsxSheetWindow:
|
||||
"""GET /api/file/{vault}/xlsx/sheet — one window of one sheet."""
|
||||
|
||||
def _get(self, client, path="budget.xlsx", **params):
|
||||
return client.get(
|
||||
f"/api/file/{VAULT}/xlsx/sheet", params={"path": path, **params}
|
||||
)
|
||||
|
||||
def test_window_returns_rows_and_totals(self, client, xlsx_file):
|
||||
resp = self._get(client, sheet="Budget", offset=0, limit=10)
|
||||
assert resp.status_code == 200
|
||||
data = resp.json()
|
||||
assert data["sheet"] == "Budget"
|
||||
assert (data["offset"], data["limit"]) == (0, 10)
|
||||
assert data["rows"] == 2 and data["cols"] == 2
|
||||
assert data["total_rows"] == 2 and data["truncated"] is False
|
||||
assert (data["max_rows"], data["max_cols"]) == (500, 40)
|
||||
assert data["has_more"] is False
|
||||
assert 'data-cell="A1"' in data["html"]
|
||||
|
||||
def test_window_keeps_the_real_a1_coordinates(self, client, xlsx_file):
|
||||
"""A window must be indistinguishable from a full render: the A1
|
||||
references and the row numbers have to be the sheet's, not the
|
||||
window's, or an edit would land on the wrong cell. The CONTENT matters
|
||||
as much as the label — row 2 of "Budget" is "Total", not "Poste"."""
|
||||
data = self._get(client, sheet="Budget", offset=1, limit=1).json()
|
||||
assert data["rows"] == 1
|
||||
assert 'data-cell="A2"' in data["html"]
|
||||
assert 'data-cell="A1"' not in data["html"]
|
||||
assert "<th class=\"xlsx-rownum\">2</th>" in data["html"]
|
||||
assert "Total" in data["html"] and "Poste" not in data["html"]
|
||||
|
||||
def test_window_offsets_walk_the_whole_sheet(self, client, wide_xlsx):
|
||||
first = self._get(client, path="grand.xlsx", sheet="Data", offset=0, limit=10).json()
|
||||
last = self._get(client, path="grand.xlsx", sheet="Data", offset=500, limit=10).json()
|
||||
assert first["truncated"] is True and first["has_more"] is True
|
||||
assert 'data-cell="A501"' in last["html"] # the row the caps used to hide
|
||||
assert "dernière ligne" in last["html"]
|
||||
assert "dernière ligne" not in first["html"]
|
||||
assert last["rows"] == 1 and last["has_more"] is False
|
||||
|
||||
def test_offset_past_the_end_is_empty_not_an_error(self, client, xlsx_file):
|
||||
resp = self._get(client, sheet="Budget", offset=9999, limit=10)
|
||||
assert resp.status_code == 200
|
||||
data = resp.json()
|
||||
assert data["rows"] == 0
|
||||
assert data["has_more"] is False
|
||||
|
||||
def test_cached_result_survives_the_window(self, client, lossy_xlsx):
|
||||
"""A12 must not be lost on the lazy path."""
|
||||
data = self._get(client, path="risky.xlsx", sheet="Data", offset=0, limit=10).json()
|
||||
assert "xlsx-cached" in data["html"]
|
||||
|
||||
def test_unknown_sheet_is_404(self, client, xlsx_file):
|
||||
resp = self._get(client, sheet="Nope")
|
||||
assert resp.status_code == 404
|
||||
assert "Feuille introuvable" in resp.json()["detail"]
|
||||
|
||||
def test_missing_file_is_404(self, client, test_vault_dir):
|
||||
assert self._get(client, path="absent.xlsx", sheet="Budget").status_code == 404
|
||||
|
||||
def test_non_xlsx_file_is_415(self, client, test_vault_dir):
|
||||
(Path(test_vault_dir) / "note.md").write_text("# hi", encoding="utf-8")
|
||||
resp = self._get(client, path="note.md", sheet="Budget")
|
||||
assert resp.status_code == 415
|
||||
|
||||
def test_limit_above_the_server_cap_is_rejected(self, client, xlsx_file):
|
||||
"""The cap is a contract, not a silent truncation of the request."""
|
||||
assert self._get(client, sheet="Budget", limit=100_000).status_code == 422
|
||||
|
||||
def test_reader_clamps_a_hostile_limit(self, xlsx_file):
|
||||
"""Belt and braces: the reader caps too, whoever calls it."""
|
||||
from backend.xlsx_reader import MAX_WINDOW_ROWS, read_sheet_window
|
||||
|
||||
window = read_sheet_window(Path(xlsx_file), "Budget", offset=0, limit=10**9)
|
||||
assert window["limit"] == MAX_WINDOW_ROWS
|
||||
|
||||
def test_reader_rejects_a_negative_offset(self, xlsx_file):
|
||||
from backend.xlsx_reader import read_sheet_window
|
||||
|
||||
window = read_sheet_window(Path(xlsx_file), "Budget", offset=-5, limit=10)
|
||||
assert window["offset"] == 0
|
||||
|
||||
|
||||
# ── #153 A4 — formula injection ──────────────────────────────────────────
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user