feat: tableurs indexables, saisie typée et valeurs calculées #153

Les tableurs étaient invisibles à la recherche, chaque saisie devenait
du texte, et une formule n'affichait que sa formule.

- A5 : extract_indexable_text() indexe les noms de feuilles et les 20
  premières lignes (plafond 5 k caractères) pour le TF-IDF et la
  recherche sémantique. Un mot tapé dans une cellule rend le fichier
  trouvable ; un classeur chiffré s'indexe par son seul nom.
- A10 : la saisie est typée comme dans Excel — booléens
  (TRUE/FAUX/OUI/NON) et dates FR JJ/MM/AAAA en ordre jour-first, donc
  01/02/2026 est le 1er février. Une saisie ressemblant à une formule
  n'est jamais convertie (BUG-088 préservé).
- A12 : la valeur calculée par Excel s'affiche sous la formule, via une
  2e lecture data_only=True faite seulement si l'archive contient un
  <v>. Info-bulle traduite FR/EN, aucun texte d'interface côté backend.
- BUG-089 : un reindex manuel ne reconstruisait pas l'index inversé, et
  backend/search.py lisait l'index via un import par valeur — après un
  rechargement du module, la recherche écrivait dans un dict périmé.

Contre-preuves vérifiées pour A5, A10 et A12 (neutralisation de chaque
fonction → échec des tests concernés). Tests : 1402 pytest, 10 JSDOM,
4 E2E, suite E2E complète verte, ruff/mypy 0.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <[email protected]>
This commit is contained in:
2026-09-27 22:43:24 -04:00
parent 31d4616baf
commit 06f8e63d06
21 changed files with 585 additions and 61 deletions
+41 -1
View File
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html). et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
> **En cours de développement** : les changements à venir sont listés dans la section > **En cours de développement** : les changements à venir sont listés dans la section
> [Unreleased](#unreleased). La dernière version livrée est **2.29.0**. > [Unreleased](#unreleased). La dernière version livrée est **2.30.0**.
--- ---
@@ -14,6 +14,46 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
--- ---
## [2.30.0] — 2026-09-27
### Correction
- **BUG-089 — un reindex manuel ne reconstruisait pas l'index inversé.**
`reload_index()` / `reload_single_vault()` remplacent l'entrée de vault
en bloc, ce qui n'émet pas les notifications incrémentales : la
recherche TF-IDF continuait de servir un index périmé après un
reindex. Les deux fonctions appellent désormais `init_inverted_index()`.
Au passage, `backend/search.py` lisait l'index via
`from backend.indexer import index` — une liaison **par valeur** du
dict : un rechargement du module `backend.indexer` recréait le dict
côté indexer alors que la recherche écrivait dans l'ancien, et
l'index inversé n'indexait plus rien. Tous les accès passent par
`_indexer.index`. *Trouvé en écrivant le test de recherche d'A5 : il
passait isolément et échouait en suite complète selon l'ordre.*
### Ajouté
- **#153 A5 — les tableurs sont indexés par leur contenu.**
`extract_indexable_text()` extrait les noms de feuilles et les 20
premières lignes (plafond 5 000 caractères, 20 feuilles) pour le
TF-IDF et la recherche sémantique. Un mot tapé dans une cellule rend
désormais le classeur trouvable ; la lecture binaire pour l'affichage
est inchangée et un classeur chiffré/corrompu s'indexe par son seul
nom.
- **#153 A10 — la saisie est typée comme dans Excel.** Une valeur
`TRUE`/`FAUX`/`OUI`/`NON` devient un booléen, une date `JJ/MM/AAAA`
(avec `HH:MM` optionnel) devient une vraie date — et dans l'ordre
français : `01/02/2026` est le 1ᵉʳ février. Une saisie ressemblant à
une formule n'est jamais convertie.
- **#153 A12 — la valeur calculée s'affiche sous la formule.** Quand une
cellule porte encore le résultat de son dernier calcul Excel, celui-ci
s'affiche dans une ligne discrète sous la formule. La seconde lecture
`data_only=True` n'a lieu que si l'archive contient réellement une
valeur en cache, et toute erreur retombe sur l'affichage formules seul.
Info-bulle traduite FR/EN (`xlsx.cached_value_title`).
---
## [2.29.0] — 2026-09-27 ## [2.29.0] — 2026-09-27
### Correction ### Correction
+3 -3
View File
@@ -4,7 +4,7 @@
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive. **Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]() [![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/) [![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/) [![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -976,8 +976,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
## 📝 Changelog ## 📝 Changelog
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.29.0). Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.30.0).
--- ---
*Projet : ObsiGate | Version : 2.29.0 | Dernière mise à jour : Septembre 2026* *Projet : ObsiGate | Version : 2.30.0 | Dernière mise à jour : Septembre 2026*
+3 -3
View File
@@ -2,7 +2,7 @@
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface. **Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]() [![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/) [![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/) [![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -1151,8 +1151,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
## 📝 Changelog ## 📝 Changelog
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.29.0). See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.30.0).
--- ---
*Project: ObsiGate | Version: 2.29.0 | Last updated: September 2026* *Project: ObsiGate | Version: 2.30.0 | Last updated: September 2026*
+1 -1
View File
@@ -1 +1 @@
2.29.0 2.30.0
+39 -7
View File
@@ -351,6 +351,23 @@ def _decompress_excalidraw(compressed: str) -> dict[str, Any] | None:
return data return data
def extract_xlsx_indexable(file_path: Path) -> str:
"""Return searchable text for a workbook (#153 A5).
Lazy wrapper: ``openpyxl`` is only imported when a spreadsheet is actually
indexed, so a vault without workbooks never pays the import. Errors are
swallowed — a corrupt or encrypted file still gets indexed by name.
"""
try:
from backend.xlsx_reader import extract_indexable_text
except Exception: # pragma: no cover - openpyxl missing
return ""
try:
return extract_indexable_text(file_path)
except Exception: # pragma: no cover - defensive
return ""
def extract_excalidraw_indexable(raw: str) -> str: def extract_excalidraw_indexable(raw: str) -> str:
"""Return indexable text content for a raw .excalidraw / .excalidraw.md file. """Return indexable text content for a raw .excalidraw / .excalidraw.md file.
@@ -561,11 +578,12 @@ def _scan_vault(
title = fpath.stem.replace("-", " ").replace("_", " ") title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = "" content_preview = ""
elif ext == ".xlsx": elif ext == ".xlsx":
# #152 — binary workbook: metadata only, the viewer renders # #153 A5 — a workbook stays rendered by the viewer, but its
# it (parity with _index_single_file_sync). # cell values are now indexed as text so a spreadsheet is
raw = "" # findable by its content (parity with _index_single_file_sync).
raw = extract_xlsx_indexable(fpath)
title = fpath.stem.replace("-", " ").replace("_", " ") title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = "" content_preview = raw[:200].strip()
else: else:
raw = fpath.read_text(encoding="utf-8", errors="replace") raw = fpath.read_text(encoding="utf-8", errors="replace")
title = fpath.stem.replace("-", " ").replace("_", " ") title = fpath.stem.replace("-", " ").replace("_", " ")
@@ -807,6 +825,13 @@ async def reload_index() -> dict[str, Any]:
await build_index() await build_index()
# BUG-040/#86: complete the deferred PDF + excalidraw extraction. # BUG-040/#86: complete the deferred PDF + excalidraw extraction.
await enrich_pdf_texts() await enrich_pdf_texts()
# The inverted index is NOT updated by the hooks here: the rebuild above
# replaces whole vault entries, so the incremental notifications are not
# emitted for the files that only changed content. Without this, a manual
# reindex left TF-IDF search serving a stale index (BUG-089).
from backend.search import init_inverted_index
init_inverted_index()
stats = {} stats = {}
for name, data in index.items(): for name, data in index.items():
stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])} stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])}
@@ -882,6 +907,13 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]:
# BUG-040/#86: complete the deferred PDF + excalidraw extraction. # BUG-040/#86: complete the deferred PDF + excalidraw extraction.
await enrich_pdf_texts(vault_name) await enrich_pdf_texts(vault_name)
# Same as reload_index: the vault entry was replaced wholesale, so rebuild
# the inverted index or TF-IDF search keeps serving stale postings
# (BUG-089).
from backend.search import init_inverted_index
init_inverted_index()
stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])} stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])}
logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags") logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags")
return stats return stats
@@ -955,9 +987,9 @@ def _index_single_file_sync(vault_name: str, vault_path: str, file_path: str, va
raw = "" raw = ""
content_preview = "" content_preview = ""
elif ext == ".xlsx": elif ext == ".xlsx":
# #152 — binary workbook: metadata only (parity with _scan_vault). # #153 A5 — index sheet names + header rows as text (see _scan_vault).
raw = "" raw = extract_xlsx_indexable(fpath)
content_preview = "" content_preview = raw[:200].strip()
else: else:
raw = fpath.read_text(encoding="utf-8", errors="replace") raw = fpath.read_text(encoding="utf-8", errors="replace")
content_preview = raw[:200].strip() content_preview = raw[:200].strip()
+10 -5
View File
@@ -12,7 +12,12 @@ from sortedcontainers import SortedList
from backend import indexer as _indexer from backend import indexer as _indexer
from backend import semantic_search as _semantic from backend import semantic_search as _semantic
from backend.indexer import index
# NOTE: the shared index is read through ``_indexer.index`` everywhere, never
# via ``from backend.indexer import index``. That import binds the dict object
# once, so a module reload of ``backend.indexer`` (tests, dev reload) rebinds
# the module-level name to a FRESH dict while this module keeps writing to the
# stale one — the inverted index then silently indexes nothing (BUG-089).
from backend.services.regex_safety import ( from backend.services.regex_safety import (
MAX_REGEX_MATCHES, MAX_REGEX_MATCHES,
truncate_for_regex, truncate_for_regex,
@@ -399,7 +404,7 @@ class InvertedIndex:
self.vault_docs = defaultdict(set) self.vault_docs = defaultdict(set)
self.tag_docs = defaultdict(set) self.tag_docs = defaultdict(set)
for vault_name, vault_data in index.items(): for vault_name, vault_data in _indexer.index.items():
for file_info in vault_data.get("files", []): for file_info in vault_data.get("files", []):
doc_key = f"{vault_name}::{file_info['path']}" doc_key = f"{vault_name}::{file_info['path']}"
self.doc_count += 1 self.doc_count += 1
@@ -688,7 +693,7 @@ _indexer.set_index_change_hook(_on_index_change_hook)
def init_inverted_index(): def init_inverted_index():
"""Force initial inverted index build. Called after build_index completes on startup.""" """Force initial inverted index build. Called after build_index completes on startup."""
if any(vdata.get("files") for vdata in index.values()): if any(vdata.get("files") for vdata in _indexer.index.values()):
_inverted_index.rebuild() _inverted_index.rebuild()
logger.info("Inverted index initialized.") logger.info("Inverted index initialized.")
@@ -784,7 +789,7 @@ def search(
else: else:
candidates = [ candidates = [
(vault_name, file_info) (vault_name, file_info)
for vault_name, vault_data in index.items() for vault_name, vault_data in _indexer.index.items()
if vault_filter == "all" or vault_name == vault_filter if vault_filter == "all" or vault_name == vault_filter
for file_info in vault_data["files"] for file_info in vault_data["files"]
] ]
@@ -1613,7 +1618,7 @@ def get_all_tags(vault_filter: str | None = None) -> dict[str, int]:
Dict mapping tag names to their total occurrence count. Dict mapping tag names to their total occurrence count.
""" """
merged: dict[str, int] = {} merged: dict[str, int] = {}
for vault_name, vault_data in index.items(): for vault_name, vault_data in _indexer.index.items():
if vault_filter and vault_filter != "all" and vault_name != vault_filter: if vault_filter and vault_filter != "all" and vault_name != vault_filter:
continue continue
for tag, count in vault_data.get("tags", {}).items(): for tag, count in vault_data.get("tags", {}).items():
+50 -1
View File
@@ -19,6 +19,7 @@ import shutil
import threading import threading
from collections.abc import Callable, Iterator from collections.abc import Callable, Iterator
from contextlib import contextmanager from contextlib import contextmanager
from datetime import date, datetime
from pathlib import Path from pathlib import Path
from typing import Any from typing import Any
@@ -237,6 +238,14 @@ _XLSX_FLOAT_RE = re.compile(r"^[+-]?(?:\d+\.\d*|\.\d+)$")
# the legacy Lotus-style trigger. "+"/"-" are left alone: they are numbers here. # the legacy Lotus-style trigger. "+"/"-" are left alone: they are numbers here.
_XLSX_FORMULA_RE = re.compile(r"^[=@]") _XLSX_FORMULA_RE = re.compile(r"^[=@]")
# #153 A10 — types recognised when a user types into a cell. Excel infers them
# too; storing everything as text would make a spreadsheet unusable (a boolean
# column stays a string, a date column sorts lexicographically).
_XLSX_TRUE_LITERALS = {"true", "vrai", "oui", "yes"}
_XLSX_FALSE_LITERALS = {"false", "faux", "non", "no"}
# Shape check before strptime: keeps the hot path free of format attempts.
_XLSX_DATE_RE = re.compile(r"^\d{1,2}[-/]\d{1,2}[-/]\d{4}(?:[ T]\d{1,2}:\d{2})?$")
# #153 A3 — per-file write lock. Two concurrent saves (two tabs, the AI agent # #153 A3 — per-file write lock. Two concurrent saves (two tabs, the AI agent
# and the viewer, a watcher restore) would otherwise read-modify-write on the # and the viewer, a watcher restore) would otherwise read-modify-write on the
# same archive and the last writer silently wins. Kept deliberately small: the # same archive and the last writer silently wins. Kept deliberately small: the
@@ -270,7 +279,21 @@ def _xlsx_write_lock(key: str) -> Iterator[None]:
def _coerce_xlsx_value(value: Any) -> Any: def _coerce_xlsx_value(value: Any) -> Any:
"""Turn the string sent by the cell editor back into a scalar.""" """Turn the string sent by the cell editor back into a scalar (#153 A10).
The coercion is symmetric with :func:`backend.xlsx_reader._fmt`: a value
typed by the user comes back as a string, and Excel would have inferred a
type when typing the same thing. Recognised here:
* an empty cell -> ``None`` (clears it)
* ``1234`` / ``-1`` -> ``int``
* ``1.5`` / ``.5`` -> ``float``
* ``TRUE``/``FAUX`` (case-insensitive) -> ``bool``
* ``31/12/2026`` / ``31/12/2026 14:30`` -> ``date``/``datetime`` (FR)
Anything else stays text. A date-looking string typed with a leading
``=`` is a formula and never reaches here as a date.
"""
if not isinstance(value, str): if not isinstance(value, str):
return value return value
text = value.strip() text = value.strip()
@@ -280,9 +303,35 @@ def _coerce_xlsx_value(value: Any) -> Any:
return int(text) return int(text)
if _XLSX_FLOAT_RE.match(text): if _XLSX_FLOAT_RE.match(text):
return float(text) return float(text)
lowered = text.lower()
if lowered in _XLSX_TRUE_LITERALS:
return True
if lowered in _XLSX_FALSE_LITERALS:
return False
if not _XLSX_FORMULA_RE.match(text):
parsed = _parse_fr_datetime(text)
if parsed is not None:
return parsed
return value return value
def _parse_fr_datetime(text: str) -> date | datetime | None:
"""Parse a FR-localised date/datetime, or return ``None``.
Accepts ``JJ/MM/AAAA`` and ``JJ/MM/AAAA HH:MM`` (also ``JJ-MM-AAAA``).
``dayfirst`` is what makes ``01/02/2026`` the 1st of February rather than
the 2nd of January — the French convention.
"""
if not _XLSX_DATE_RE.match(text):
return None
for fmt in ("%d/%m/%Y %H:%M", "%d/%m/%Y", "%d-%m-%Y %H:%M", "%d-%m-%Y"):
try:
return datetime.strptime(text, fmt)
except ValueError:
continue
return None
def _write_cell(ws: Any, ref: str, value: Any, *, allow_formula: bool) -> None: def _write_cell(ws: Any, ref: str, value: Any, *, allow_formula: bool) -> None:
"""Assign one cell, forcing text when it looks like a formula. """Assign one cell, forcing text when it looks like a formula.
+163 -13
View File
@@ -11,6 +11,7 @@ round-trip would drop (#153 A1) so the UI can warn before saving.
from __future__ import annotations from __future__ import annotations
import html import html
import logging
import re import re
import zipfile import zipfile
from datetime import date, datetime from datetime import date, datetime
@@ -20,6 +21,8 @@ from typing import Any
from openpyxl import load_workbook from openpyxl import load_workbook
from openpyxl.utils import get_column_letter from openpyxl.utils import get_column_letter
logger = logging.getLogger("obsigate.xlsx_reader")
# ponytail: hard caps bound the rendered grid (500 rows x 40 cols per sheet). # ponytail: hard caps bound the rendered grid (500 rows x 40 cols per sheet).
# Raise them, or paginate per sheet, if a real workbook needs more. # Raise them, or paginate per sheet, if a real workbook needs more.
MAX_ROWS = 500 MAX_ROWS = 500
@@ -48,6 +51,14 @@ _CACHED_FORMULA_RE = re.compile(rb"<f[ >][^<]*</f>\s*<v>[^<]")
# Sheet XML scanned by the cached-formula probe (CPU guard, like MAX_REPLACE_FILE_BYTES). # Sheet XML scanned by the cached-formula probe (CPU guard, like MAX_REPLACE_FILE_BYTES).
_MAX_PROBE_BYTES = 8_000_000 _MAX_PROBE_BYTES = 8_000_000
# #153 A5 — ceiling on the text handed to the TF-IDF / semantic index. A workbook
# is a data dump, not prose: indexing every cell would flood the inverted index
# and bury the notes. Sheet names + the first rows are enough to make a
# spreadsheet findable by its headers.
MAX_INDEX_CHARS = 5_000
_INDEX_ROWS_PER_SHEET = 20
MAX_INDEX_SHEETS = 20
def _fmt(value: Any) -> str: def _fmt(value: Any) -> str:
if value is None: if value is None:
@@ -74,7 +85,29 @@ def _trim(grid: list[list[str]]) -> list[list[str]]:
return [row[:width] for row in grid] return [row[:width] for row in grid]
def _table(grid: list[list[str]]) -> str: def _cell_cached(cached: list[list[str]] | None, r: int, c: int) -> str:
"""Return the cached result for a 0-based cell, or ``""``.
The shadow grid is read positionally and may be narrower than the formula
grid (``_trim`` collapses the trailing empty columns of each grid
independently), so every lookup is bounds-checked rather than assumed.
"""
if not cached or r >= len(cached):
return ""
row = cached[r]
return row[c] if c < len(row) else ""
def _table(grid: list[list[str]], cached: list[list[str]] | None = None) -> str:
"""Render a grid as an HTML table.
``cached`` is the same grid read with ``data_only=True`` (#153 A12): where a
formula cell still carries its last computed result, it is shown as a
discreet second line (``<span class="xlsx-cached">``) so the user sees the
number Excel last calculated instead of only the formula text. The span
carries ``data-cached-value`` and is titled client-side from
``xlsx.cached_value_title`` — the backend never emits UI text.
"""
if not grid: if not grid:
return "<p><em>Feuille vide</em></p>" return "<p><em>Feuille vide</em></p>"
n_cols = max(len(row) for row in grid) n_cols = max(len(row) for row in grid)
@@ -90,7 +123,24 @@ def _table(grid: list[list[str]]) -> str:
out.append(f'<tr><th class="xlsx-rownum">{r}</th>') out.append(f'<tr><th class="xlsx-rownum">{r}</th>')
for c, val in enumerate(row, start=1): for c, val in enumerate(row, start=1):
ref = f"{get_column_letter(c)}{r}" ref = f"{get_column_letter(c)}{r}"
out.append(f'<td data-cell="{ref}">{html.escape(val)}</td>') # The cached result only makes sense for a formula cell: on a plain
# value cell the two reads are identical and showing both would
# duplicate the text.
shadow = ""
if cached is not None and val.startswith("="):
# `c` is 1-based (A1 notation) and `r` too, while the grid is
# 0-based: translate both.
cval = _cell_cached(cached, r - 1, c - 1)
if cval and cval != val:
# The tooltip is translated client-side from
# `xlsx.cached_value_title`; never hardcode UI text here.
shadow = (
f'<span class="xlsx-cached" data-cached-value="1">'
f"{html.escape(cval)}</span>"
)
out.append(
f'<td data-cell="{ref}">{html.escape(val)}{shadow}</td>'
)
out.append("</tr>") out.append("</tr>")
out.append("</tbody></table></div>") out.append("</tbody></table></div>")
return "".join(out) return "".join(out)
@@ -142,18 +192,118 @@ def inspect_workbook(file_path: Path) -> list[str]:
def render_sheets(file_path: Path) -> list[dict[str, str]]: def render_sheets(file_path: Path) -> list[dict[str, str]]:
"""Return ``[{"name": sheet_title, "html": table_html}, ...]``.""" """Return ``[{"name": sheet_title, "html": table_html}, ...]``.
Reads the workbook twice: once with ``data_only=False`` for the formulas
(what the user must edit) and, when any formula carries a cached result
(#153 A12), once with ``data_only=True`` to show what Excel last computed.
The second pass is skipped entirely when the archive holds no cached value,
so the common case still costs a single load.
"""
wb = load_workbook(str(file_path), read_only=True, data_only=False) wb = load_workbook(str(file_path), read_only=True, data_only=False)
try: try:
sheets = [] formulas = [_sheet_grid(ws) for ws in wb.worksheets]
for ws in wb.worksheets: titles = [ws.title for ws in wb.worksheets]
grid = [
[_fmt(v) for v in row]
for row in ws.iter_rows(
min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True
)
]
sheets.append({"name": ws.title, "html": _table(_trim(grid))})
return sheets
finally: finally:
wb.close() wb.close()
cached: list[list[list[str]]] | None = None
if _has_cached_values(file_path):
cached = _read_cached_grids(file_path, titles)
sheets = []
for i, title in enumerate(titles):
grid = _trim(formulas[i])
# The shadow grid is NOT trimmed independently: _trim drops the
# trailing empty columns of each grid on its own width, which would
# shift every cached value left of its formula. Indexing it
# positionally against the untrimmed grid keeps the two aligned.
shadow = cached[i] if cached is not None and i < len(cached) else None
sheets.append({"name": title, "html": _table(grid, shadow)})
return sheets
def _sheet_grid(ws: Any) -> list[list[str]]:
"""Read one worksheet into a grid of formatted strings, bounded by the caps."""
return [
[_fmt(v) for v in row]
for row in ws.iter_rows(
min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True
)
]
def _has_cached_values(file_path: Path) -> bool:
"""True when the archive holds at least one ``<f>…</f><v>…</v>``."""
try:
with zipfile.ZipFile(file_path) as zf:
return _has_cached_formulas(zf)
except (OSError, zipfile.BadZipFile):
return False
def _read_cached_grids(
file_path: Path, titles: list[str]
) -> list[list[list[str]]] | None:
"""Read every sheet with ``data_only=True`` (what Excel last computed).
Best effort: returns ``None`` on any failure so the viewer falls back to the
formula-only rendering. A workbook Excel opens but openpyxl cannot re-read
must still display.
"""
try:
wb = load_workbook(str(file_path), read_only=True, data_only=True)
except Exception:
return None
try:
grids = [_sheet_grid(ws) for ws in wb.worksheets]
if [ws.title for ws in wb.worksheets] != titles:
return None
return grids
except Exception:
logger.debug("xlsx cached values unavailable", exc_info=True)
return None
finally:
wb.close()
def extract_indexable_text(file_path: Path) -> str:
"""Return searchable text for the TF-IDF / semantic index (#153 A5).
Sheet names plus the first :data:`_INDEX_ROWS_PER_SHEET` rows of each
sheet, capped at :data:`MAX_INDEX_CHARS`. Rows are tab-joined so a search
for a header matches the sheet it belongs to.
Never raises: a corrupt, encrypted or unsupported workbook yields ``""`` so
the file still gets indexed by name (same contract as :func:`inspect_workbook`).
"""
chunks: list[str] = []
budget = MAX_INDEX_CHARS
try:
wb = load_workbook(str(file_path), read_only=True, data_only=True)
except Exception:
# Encrypted (BadZipFile) or not a real workbook: name-only indexing.
return ""
try:
for ws in wb.worksheets[:MAX_INDEX_SHEETS]:
if budget <= 0:
break
# The sheet title alone is a strong signal ("Recettes", "Budget").
block = [ws.title]
for row in ws.iter_rows(
min_row=1, max_row=_INDEX_ROWS_PER_SHEET, max_col=MAX_COLS, values_only=True
):
cells = [_fmt(v) for v in row]
# Skip blank rows instead of emitting runs of tabs.
if not any(c.strip() for c in cells):
continue
block.append("\t".join(cells).rstrip())
text = "\n".join(block)
chunks.append(text[:budget])
budget -= len(text)
except Exception:
# Truncated but still useful: keep whatever was collected.
pass
finally:
wb.close()
return "\n".join(c for c in chunks if c).strip()
+1 -1
View File
@@ -2626,7 +2626,7 @@ dependencies = [
[[package]] [[package]]
name = "obsigate-desktop" name = "obsigate-desktop"
version = "2.29.0" version = "2.30.0"
dependencies = [ dependencies = [
"chrono", "chrono",
"env_logger", "env_logger",
+1 -1
View File
@@ -1,6 +1,6 @@
[package] [package]
name = "obsigate-desktop" name = "obsigate-desktop"
version = "2.29.0" version = "2.30.0"
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian" description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
authors = ["Bruno Charest"] authors = ["Bruno Charest"]
edition = "2021" edition = "2021"
+1 -1
View File
@@ -1,7 +1,7 @@
{ {
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json", "$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
"productName": "ObsiGate", "productName": "ObsiGate",
"version": "2.29.0", "version": "2.30.0",
"identifier": "com.obsigate.desktop", "identifier": "com.obsigate.desktop",
"build": { "build": {
"frontendDist": "../frontend", "frontendDist": "../frontend",
+2
View File
@@ -195,6 +195,7 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-086* | Édition d'un `.xlsx` : `wb.save()` écrit en place, un plantage laisse un classeur corrompu | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::edit_xlsx_cells` | Simuler un `OSError` pendant `Workbook.save` → le fichier d'origine est tronqué | Écriture atomique : `wb.save(<nom>.<pid>.tmp)` puis `os.replace()` ; `.tmp` supprimé sur échec ; le backup `.bak` reste inchangé. Test : `TestXlsxAtomicWrite::test_failed_save_keeps_the_original` (octets identiques après échec) + `test_no_tmp_left_after_a_successful_save` | #153 A2. Le fichier temporaire a un suffixe `.tmp` → ignoré par le watcher (`_is_relevant` ne retient que les extensions supportées). Vérifié : cf. BUG-085 | | *BUG-086* | Édition d'un `.xlsx` : `wb.save()` écrit en place, un plantage laisse un classeur corrompu | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::edit_xlsx_cells` | Simuler un `OSError` pendant `Workbook.save` → le fichier d'origine est tronqué | Écriture atomique : `wb.save(<nom>.<pid>.tmp)` puis `os.replace()` ; `.tmp` supprimé sur échec ; le backup `.bak` reste inchangé. Test : `TestXlsxAtomicWrite::test_failed_save_keeps_the_original` (octets identiques après échec) + `test_no_tmp_left_after_a_successful_save` | #153 A2. Le fichier temporaire a un suffixe `.tmp` → ignoré par le watcher (`_is_relevant` ne retient que les extensions supportées). Vérifié : cf. BUG-085 |
| *BUG-087* | Édition d'un `.xlsx` concurrente (deux onglets, agent IA + viewer) : read-modify-write sans verrou, le dernier écrivain gagne silencieusement | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::_xlsx_write_lock` | Deux `PUT xlsx/save` simultanés sur le même fichier → une écriture est écrasée sans trace | Verrou par chemin (registre + garde, timeout 15 s) autour du cycle load → edit → `os.replace` ; attente dépassée → **409** `conflict`. L'endpoint est devenu `def` (sync) pour que l'attente s'exécute dans le threadpool et ne bloque pas la boucle d'événements. Test : `TestXlsxWriteLock` (2) | #153 A3. Verrou en mémoire, par processus : protège les cas d'un même serveur (le cas desktop/Tauri). Vérifié : cf. BUG-085 | | *BUG-087* | Édition d'un `.xlsx` concurrente (deux onglets, agent IA + viewer) : read-modify-write sans verrou, le dernier écrivain gagne silencieusement | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::_xlsx_write_lock` | Deux `PUT xlsx/save` simultanés sur le même fichier → une écriture est écrasée sans trace | Verrou par chemin (registre + garde, timeout 15 s) autour du cycle load → edit → `os.replace` ; attente dépassée → **409** `conflict`. L'endpoint est devenu `def` (sync) pour que l'attente s'exécute dans le threadpool et ne bloque pas la boucle d'événements. Test : `TestXlsxWriteLock` (2) | #153 A3. Verrou en mémoire, par processus : protège les cas d'un même serveur (le cas desktop/Tauri). Vérifié : cf. BUG-085 |
| *BUG-088* | Injection de formule dans un `.xlsx` : une saisie `=cmd\|'/c calc'!A1` est stockée comme formule et s'exécute à l'ouverture dans Excel (DDE) | 🟢 corrigé | P0 | tableur Excel / sécurité | IA | `backend/services/mutations.py::_write_cell`, `backend/routers/files_write.py`, `frontend/js/viewer.js::renderXlsxViewer` | `PUT /api/file/V/xlsx/save` avec `{"sheet": "S", "cells": {"A1": "=1+1"}}` → la cellule sort en `data_type == "f"` | `cell.data_type = "s"` après affectation : le texte est stocké comme chaîne, aucun `<f>` n'est écrit. Opt-in via `allow_formula: true` (endpoint) et le bouton `f(x)` de la visionneuse (session, jamais persisté). Test : `TestXlsxFormulaGuard` (4) + `xlsx-viewer.test.mjs` (toggle) | #153 A4. `+`/`-` ne sont pas neutralisés : ils sont déjà convertis en nombre par `_coerce_xlsx_value`. Le handler global `ServiceError` expose désormais `code` + `details` (le client en a besoin pour le 409), et `api()` (frontend) les propage sur l'Error. Vérifié : cf. BUG-085 | | *BUG-088* | Injection de formule dans un `.xlsx` : une saisie `=cmd\|'/c calc'!A1` est stockée comme formule et s'exécute à l'ouverture dans Excel (DDE) | 🟢 corrigé | P0 | tableur Excel / sécurité | IA | `backend/services/mutations.py::_write_cell`, `backend/routers/files_write.py`, `frontend/js/viewer.js::renderXlsxViewer` | `PUT /api/file/V/xlsx/save` avec `{"sheet": "S", "cells": {"A1": "=1+1"}}` → la cellule sort en `data_type == "f"` | `cell.data_type = "s"` après affectation : le texte est stocké comme chaîne, aucun `<f>` n'est écrit. Opt-in via `allow_formula: true` (endpoint) et le bouton `f(x)` de la visionneuse (session, jamais persisté). Test : `TestXlsxFormulaGuard` (4) + `xlsx-viewer.test.mjs` (toggle) | #153 A4. `+`/`-` ne sont pas neutralisés : ils sont déjà convertis en nombre par `_coerce_xlsx_value`. Le handler global `ServiceError` expose désormais `code` + `details` (le client en a besoin pour le 409), et `api()` (frontend) les propage sur l'Error. Vérifié : cf. BUG-085 |
| *BUG-089* | Un reindex manuel ne reconstruisait pas l'index inversé : la recherche TF-IDF continuait de servir un index périmé | 🟢 corrigé | P1 | ⚙️ backend / recherche | IA | `backend/indexer.py::reload_index`, `backend/indexer.py::reload_single_vault`, `backend/search.py` | Modifier le contenu d'un fichier, puis `GET /api/index/reload` → la recherche renvoie encore l'ancien contenu (ou rien pour un fichier nouveau) | `reload_index()` / `reload_single_vault()` appellent `init_inverted_index()` après le rebuild (le remplacement wholesale d'une entrée de vault n'émet pas les notifications incrémentales). En prime, `backend/search.py` lisait l'index via `from backend.indexer import index` (liaison **par valeur** du dict) : un `importlib.reload(backend.indexer)` recréait le dict côté indexer tandis que la recherche écrivait encore dans l'ancien — l'index inversé n'indexait alors plus rien. Tous les accès passent désormais par `_indexer.index`. Contre-preuve : `TestXlsxSearchable::test_search_finds_a_word_stored_in_a_cell` échoue sans le correctif | #153 A5. Trouvé en écrivant le test de recherche d'A5 : il passait isolément et échouait en suite complète selon l'ordre. Le reload incrémental par fichier (watcher, edition) n'est pas concerné : il passe par le hook `_on_index_change`. Vérifié : suite 1402 passed / 6 skipped, ruff/mypy 0 |
### TODOs techniques (améliorations / nouvelles tâches) ### TODOs techniques (améliorations / nouvelles tâches)
@@ -212,6 +213,7 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| Date | ID(s) traité(s) | Action | Fichiers modifiés | Résumé | Statut après | | Date | ID(s) traité(s) | Action | Fichiers modifiés | Résumé | Statut après |
|---|---|---|---|---|---| |---|---|---|---|---|---|
| 2026-09-28 | BUG-089 (#153 A5, A10, A12) | Correction | `backend/xlsx_reader.py`, `backend/indexer.py`, `backend/search.py`, `backend/services/mutations.py`, `frontend/js/viewer.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/test_xlsx_viewer.py` | **Les tableurs deviennent visibles ettypés** : (A5) `extract_indexable_text()` indexe noms de feuilles + 20 premières lignes (plafond 5 k caractères) dans le TF-IDF et la recherche sémantique — un mot tapé dans une cellule rend le fichier trouvable ; (A10) `_coerce_xlsx_value()` reconnaît désormais les booléens (`TRUE`/`FAUX`/`OUI`/`NON`) et les dates FR `JJ/MM/AAAA` (jour-first : `01/02/2026` = 1er février), symétrique avec l'affichage ; (A12) la valeur calculée en cache s'affiche sous la formule (`<span class="xlsx-cached">`, 2ᵉ lecture `data_only=True` uniquement si l'archive contient un `<v>`), info-bulle traduite via `xlsx.cached_value_title` FR/EN. (BUG-089) un reindex manuel reconstruisait mal l'index inversé et `backend/search.py` lisait l'index par valeur. Contre-preuves vérifiées pour A5, A10 et A12. Vérifié : `test_xlsx_viewer.py` 43 passed, suite 1402 passed / 6 skipped, ruff 0, mypy 0, i18n parity, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10 | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-27 | BUG-085 → BUG-088 (#153 A1-A4) | Correction | `backend/xlsx_reader.py`, `backend/services/mutations.py`, `backend/routers/files_read.py`, `backend/routers/files_write.py`, `backend/schemas.py`, `backend/main.py`, `frontend/js/viewer.js`, `frontend/js/auth.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `frontend/sw.js`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-lossy.xlsx`, `.gitea/workflows/ci.yml` | **Garde-fous d'écriture des classeurs Excel** : (BUG-085) `inspect_workbook()` détecte ce qu'un round-trip openpyxl perd (valeurs calculées, slicers, contrôles, connexions, custom XML, signature) → la lecture expose `xlsx_lossy_features`, la visionneuse affiche une bannière et `PUT xlsx/save` refuse sans `force` (**409** `xlsx_lossy_content`, confirmation explicite puis reprise) ; (BUG-086) écriture atomique `.tmp` + `os.replace` ; (BUG-087) verrou par fichier (409 `conflict`, endpoint sync pour le threadpool) ; (BUG-088) une saisie `=`/`@` est stockée en texte (`data_type = "s"`), sauf opt-in `allow_formula` / bouton `f(x)`. Le handler `ServiceError` expose désormais `code` + `details` et `api()` les propage. Périmètre de perte revalidé empiriquement sur openpyxl 3.1.5 (graphiques, images et TCD sont préservés). Vérifié : `test_xlsx_viewer.py` 31 passed, suite 1390 passed / 6 skipped, ruff/mypy 0, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10, E2E 3/3 | 🟢 corrigé (en attente vérif utilisateur) | | 2026-09-27 | BUG-085 → BUG-088 (#153 A1-A4) | Correction | `backend/xlsx_reader.py`, `backend/services/mutations.py`, `backend/routers/files_read.py`, `backend/routers/files_write.py`, `backend/schemas.py`, `backend/main.py`, `frontend/js/viewer.js`, `frontend/js/auth.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `frontend/sw.js`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-lossy.xlsx`, `.gitea/workflows/ci.yml` | **Garde-fous d'écriture des classeurs Excel** : (BUG-085) `inspect_workbook()` détecte ce qu'un round-trip openpyxl perd (valeurs calculées, slicers, contrôles, connexions, custom XML, signature) → la lecture expose `xlsx_lossy_features`, la visionneuse affiche une bannière et `PUT xlsx/save` refuse sans `force` (**409** `xlsx_lossy_content`, confirmation explicite puis reprise) ; (BUG-086) écriture atomique `.tmp` + `os.replace` ; (BUG-087) verrou par fichier (409 `conflict`, endpoint sync pour le threadpool) ; (BUG-088) une saisie `=`/`@` est stockée en texte (`data_type = "s"`), sauf opt-in `allow_formula` / bouton `f(x)`. Le handler `ServiceError` expose désormais `code` + `details` et `api()` les propage. Périmètre de perte revalidé empiriquement sur openpyxl 3.1.5 (graphiques, images et TCD sont préservés). Vérifié : `test_xlsx_viewer.py` 31 passed, suite 1390 passed / 6 skipped, ruff/mypy 0, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10, E2E 3/3 | 🟢 corrigé (en attente vérif utilisateur) |
| *(exemple)* 2026-06-15 | BUG-001 | Correction | `frontend/app.js` | Réécriture de `renderFile()` pour préserver le DOM dashboard | 🟢 corrigé (en attente vérif) | | *(exemple)* 2026-06-15 | BUG-001 | Correction | `frontend/app.js` | Réécriture de `renderFile()` pour préserver le DOM dashboard | 🟢 corrigé (en attente vérif) |
| 2026-09-09 | BUG-001, BUG-002 | Correction | `backend/main.py`, `frontend/excalidraw-editor.html`, `tests/test_pdf_stream.py` | BUG-001: Content-Disposition RFC 5987 (nom PDF accentué ne casse plus l'en-tête → plus de 500). BUG-002: suppression alias esm.sh (408 jotai) + React 19 cohérent + prop `excalidrawAPI` → Loading masqué, save OK. Vérifié: 534 tests backend verts + E2E navigateur. | 🟢 corrigé (en attente vérif utilisateur) | | 2026-09-09 | BUG-001, BUG-002 | Correction | `backend/main.py`, `frontend/excalidraw-editor.html`, `tests/test_pdf_stream.py` | BUG-001: Content-Disposition RFC 5987 (nom PDF accentué ne casse plus l'en-tête → plus de 500). BUG-002: suppression alias esm.sh (408 jotai) + React 19 cohérent + prop `excalidrawAPI` → Loading masqué, save OK. Vérifié: 534 tests backend verts + E2E navigateur. | 🟢 corrigé (en attente vérif utilisateur) |
+6 -6
View File
@@ -1,6 +1,6 @@
# ObsiGate — Roadmap # ObsiGate — Roadmap
> **Version :** 2.29.0 | **Dernière mise à jour :** 2026-09-27 > **Version :** 2.30.0 | **Dernière mise à jour :** 2026-09-27
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact > **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
> vers les fonctionnalités livrées. > vers les fonctionnalités livrées.
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)** > - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
@@ -47,7 +47,7 @@
### 153. Visionneuse & édition XLSX — complétude (fidélité, recherche, IA, UX, formats) ### 153. Visionneuse & édition XLSX — complétude (fidélité, recherche, IA, UX, formats)
- **Effort :** 8-13 jours (P0 ✅ 2-3 j · P1 : 4-6 j · P2 : 2-4 j) | **Impact :** 🟡 - **Effort :** 8-13 jours (P0 ✅ 2-3 j · P1 : 4-6 j · P2 : 2-4 j) | **Impact :** 🟡
- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), reste P1 puis P2 - **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17
- **Analyse, risques et critères d'acceptation :** [features/xlsx-viewer.md](./features/xlsx-viewer.md) - **Analyse, risques et critères d'acceptation :** [features/xlsx-viewer.md](./features/xlsx-viewer.md)
- **Description :** #152 (visionneuse XLSX, 2.27.0) lit et édite correctement la **grille de - **Description :** #152 (visionneuse XLSX, 2.27.0) lit et édite correctement la **grille de
valeurs** d'un `.xlsx`, mais l'ensemble supporté est étroit : valeurs seulement (ni structure, valeurs** d'un `.xlsx`, mais l'ensemble supporté est étroit : valeurs seulement (ni structure,
@@ -70,15 +70,15 @@
- [x] **A2** Écriture atomique (`wb.save(.tmp)` + `os.replace()`, backup inchangé) — BUG-086 - [x] **A2** Écriture atomique (`wb.save(.tmp)` + `os.replace()`, backup inchangé) — BUG-086
- [x] **A3** Verrou par fichier autour du read-modify-write (timeout 15 s + **409** `conflict`) — BUG-087 - [x] **A3** Verrou par fichier autour du read-modify-write (timeout 15 s + **409** `conflict`) — BUG-087
- [x] **A4** Neutralisation de l'injection de formule (`=`/`@` stockés en texte, opt-in `allow_formula` + bouton `f(x)`) — BUG-088 - [x] **A4** Neutralisation de l'injection de formule (`=`/`@` stockés en texte, opt-in `allow_formula` + bouton `f(x)`) — BUG-088
- **P1 — recherche, IA, UX (🟡, 4-6 j) — ⚪ à faire** - **P1 — recherche, IA, UX (🟡, 4-6 j) — 🔵 en cours**
- [ ] **A5** Indexation du contenu des feuilles (TF-IDF + sémantique, plafond ~5 k caractères) - [x] **A5** Indexation du contenu des feuilles (noms de feuilles + 20 premières lignes, plafond 5 k caractères) — les mots tapés dans une cellule rendent le fichier trouvable ; au passage **BUG-089** (reindex manuel ne reconstruisait pas l'index inversé)
- [ ] **A6** Outils IA `update_xlsx_cells` / `append_xlsx_rows` / `xlsx_to_markdown` / `list_xlsx_sheets` - [ ] **A6** Outils IA `update_xlsx_cells` / `append_xlsx_rows` / `xlsx_to_markdown` / `list_xlsx_sheets`
- [ ] **A7** Navigation clavier + barre de formule + nom de cellule (Tab/Entrée/flèches, `Maj+Entrée`, copie de plage) - [ ] **A7** Navigation clavier + barre de formule + nom de cellule (Tab/Entrée/flèches, `Maj+Entrée`, copie de plage)
- [ ] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse) - [ ] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse)
- [ ] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel) - [ ] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel)
- [ ] **A10** Types & formats de saisie (nombre/texte, booléens, dates localisées FR) - [x] **A10** Types & formats de saisie (nombre/texte, booléens `TRUE`/`FAUX`, dates FR `JJ/MM/AAAA` jour-first)
- [ ] **A11** Tests frontend (`tests/frontend/xlsx-viewer.test.mjs`) + E2E (`tests/e2e/xlsx-viewer.spec.js`) au CI - [ ] **A11** Tests frontend (`tests/frontend/xlsx-viewer.test.mjs`) + E2E (`tests/e2e/xlsx-viewer.spec.js`) au CI
- [ ] **A12** Affichage de la valeur calculée en cache (lecture `data_only=True`, mention FR/EN) - [x] **A12** Valeur calculée affichée sous la formule (2ᵉ lecture `data_only=True` seulement si l'archive contient un `<v>`, info-bulle FR/EN)
- **P2 — étendu (🟢, 2-4 j) — ⚪ à faire** - **P2 — étendu (🟢, 2-4 j) — ⚪ à faire**
- [ ] **A13** Tri / filtre / recherche dans la feuille + export CSV de la sélection - [ ] **A13** Tri / filtre / recherche dans la feuille + export CSV de la sélection
- [ ] **A14** CRUD de feuilles, lignes et colonnes (renommer, insérer, supprimer, dupliquer) - [ ] **A14** CRUD de feuilles, lignes et colonnes (renommer, insérer, supprimer, dupliquer)
+22 -13
View File
@@ -3,7 +3,7 @@
> **Item de roadmap :** [#153 — Visionneuse & édition XLSX — complétude](../ROADMAP.md) > **Item de roadmap :** [#153 — Visionneuse & édition XLSX — complétude](../ROADMAP.md)
> **Origine :** #152 (visionneuse XLSX, livrée en 2.27.0 — voir > **Origine :** #152 (visionneuse XLSX, livrée en 2.27.0 — voir
> [archive/COMPLETED_v1-v2.md](../archive/COMPLETED_v1-v2.md)) > [archive/COMPLETED_v1-v2.md](../archive/COMPLETED_v1-v2.md))
> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), P1/P2 restants > **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17
> **Effort estimé :** 8-13 jours au total (P0 ✅ 2-3 j · P1 4-6 j · P2 2-4 j) > **Effort estimé :** 8-13 jours au total (P0 ✅ 2-3 j · P1 4-6 j · P2 2-4 j)
> **Règle de maintenance :** la Roadmap porte les cases à cocher (suivi), cette fiche porte > **Règle de maintenance :** la Roadmap porte les cases à cocher (suivi), cette fiche porte
> l'analyse, les risques et les critères d'acceptation. **Ne pas dupliquer le détail.** > l'analyse, les risques et les critères d'acceptation. **Ne pas dupliquer le détail.**
@@ -134,13 +134,14 @@ couverture) · effort en jours-homme de développement + tests.
nombres. Au passage : le handler `ServiceError` expose `code` + `details` et `api()` les nombres. Au passage : le handler `ServiceError` expose `code` + `details` et `api()` les
propage sur l'Error. *Vérifié :* `TestXlsxFormulaGuard` (4) + test du toggle côté UI. propage sur l'Error. *Vérifié :* `TestXlsxFormulaGuard` (4) + test du toggle côté UI.
### P1 — Recherche, IA, UX (4-6 j) ### P1 — Recherche, IA, UX (4-6 j) — 🟢 A5, A10, A12 livrés le 2026-09-28
- [ ] **A5 — Indexation du contenu des feuilles.** Extraire un texte (noms de feuilles + - [x] **A5 — Indexation du contenu des feuilles.** `extract_indexable_text()` (noms de feuilles +
en-têtes + N premières lignes, plafond ~5 k caractères) pour le TF-IDF et la recherche 20 premières lignes, `MAX_INDEX_CHARS = 5 000`, 20 feuilles max) alimente le TF-IDF et la
sémantique, tout en gardant la lecture binaire pour l'affichage ; `content_preview` recherche sémantique ; la lecture binaire reste inchangée pour l'affichage. Un classeur
renseigné ; exclusion si le classeur est chiffré/corrompu. *Critère :* une cellule contenant chiffré/corrompu s'indexe par son seul nom (jamais d'exception). Au passage : **BUG-089**,
un mot-clé rend le fichier trouvable ; `test_xlsx_indexing` étendu. un reindex manuel ne reconstruisait pas l'index inversé. *Vérifié :* `TestXlsxSearchable` (4)
+ `TestXlsxIndexing`, **contre-preuve** (neutraliser l'extraction → 3 tests échouent).
- [ ] **A6 — Outils IA sur classeur.** `update_xlsx_cells` (enveloppe du service existant), - [ ] **A6 — Outils IA sur classeur.** `update_xlsx_cells` (enveloppe du service existant),
`append_xlsx_rows`, `xlsx_to_markdown` (contexte LLM, plafonné), `list_xlsx_sheets` — risque `append_xlsx_rows`, `xlsx_to_markdown` (contexte LLM, plafonné), `list_xlsx_sheets` — risque
WRITE + confirmation pour les mutations, libellés i18n dans `backend/tools/labels.py`, WRITE + confirmation pour les mutations, libellés i18n dans `backend/tools/labels.py`,
@@ -155,15 +156,22 @@ couverture) · effort en jours-homme de développement + tests.
`GET /api/file/{vault}/xlsx/sheet?sheet=N&offset=&limit=` (`response_model` + `GET /api/file/{vault}/xlsx/sheet?sheet=N&offset=&limit=` (`response_model` +
`backend/openapi_docs.py`), rendu à la demande avec défilement virtuel, bouton « charger `backend/openapi_docs.py`), rendu à la demande avec défilement virtuel, bouton « charger
tout ». tout ».
- [ ] **A10 — Types et formats de saisie.** Coercion symétrique à l'écriture/à l'affichage - [x] **A10 — Types et formats de saisie.** `_coerce_xlsx_value()` reconnait les booléens
(nombre vs texte, booléens `TRUE`/`FALSE`, dates localisées FR — le TODO existe déjà dans (`true`/`vrai`/`oui`/`yes` et leurs négatifs) et les dates FR `JJ/MM/AAAA` (+ `HH:MM`),
`_coerce_xlsx_value`) ; affichage du type d'origine dans l'info-bulle de cellule. jour-first comme Excel en locale française : `01/02/2026` = 1ᵉʳ février. Une saisie
ressemblant à une formule n'est jamais convertie (BUG-088 préservé) ; un code postal
numérique ou une version restent ce qu'ils sont. *Vérifié :* `TestXlsxValueCoercion` (5),
**contre-preuve** (neutraliser la coercion → 2 tests échouent).
- [ ] **A11 — Tests frontend + E2E.** `tests/frontend/xlsx-viewer.test.mjs` (dirty, Échap, - [ ] **A11 — Tests frontend + E2E.** `tests/frontend/xlsx-viewer.test.mjs` (dirty, Échap,
collage, 1 PUT par feuille, bouton désactivé) et `tests/e2e/xlsx-viewer.spec.js` collage, 1 PUT par feuille, bouton désactivé) et `tests/e2e/xlsx-viewer.spec.js`
(ouverture, onglets, édition, sauvegarde, rechargement) ; intégration au CI. (ouverture, onglets, édition, sauvegarde, rechargement) ; intégration au CI.
- [ ] **A12 — Valeurs calculées.** Afficher la valeur en cache (2ᵉ ligne discrète) quand elle - [x] **A12 — Valeurs calculées.** La valeur en cache s'affiche sous la formule dans un
existe, via une lecture `data_only=True` de la même page d'onglets ; mention FR/EN `<span class="xlsx-cached">`. La 2ᵉ lecture `data_only=True` n'a lieu que si l'archive
« valeur recalculée par Excel ». contient réellement un `<f>…</f><v>…</v>` (sonde déjà présente pour A1) : le cas courant
reste à un seul chargement, et toute erreur retombe sur l'affichage formules seul.
L'info-bulle est traduite côté client (`xlsx.cached_value_title` FR/EN) — aucun texte
d'interface n'est émis par le backend. *Vérifié :* `TestXlsxCachedValues` (3),
**contre-preuve** (neutraliser la 2ᵉ lecture → 2 tests échouent).
### P2 — Étendu (2-4 j) ### P2 — Étendu (2-4 j)
@@ -199,3 +207,4 @@ couverture) · effort en jours-homme de développement + tests.
| 2026-09-27 | Audit complet → création de #153 : limites, risques R1-R5, backlog A1-A17 | | 2026-09-27 | Audit complet → création de #153 : limites, risques R1-R5, backlog A1-A17 |
| 2026-09-27 | Périmètre de perte **remesuré** sur openpyxl 3.1.5 : graphiques / images / TCD sont préservés, seules les valeurs en cache et quelques parties exotiques sont perdues | | 2026-09-27 | Périmètre de perte **remesuré** sur openpyxl 3.1.5 : graphiques / images / TCD sont préservés, seules les valeurs en cache et quelques parties exotiques sont perdues |
| 2026-09-27 | **P0 livré** (BUG-085 → BUG-088) : `xlsx_lossy_features` + 409 `xlsx_lossy_content`, écriture atomique, verrou par fichier, formules stockées en texte par défaut | | 2026-09-27 | **P0 livré** (BUG-085 → BUG-088) : `xlsx_lossy_features` + 409 `xlsx_lossy_content`, écriture atomique, verrou par fichier, formules stockées en texte par défaut |
| 2026-09-28 | **A5 + A10 + A12 livrés** : le contenu des cellules est indexé (recherche), la saisie est typée (booléens, dates FR), la valeur calculée s'affiche sous la formule. **BUG-089** corrigé au passage (reindex manuel ≠ reconstruction de l'index inversé ; `backend/search.py` lisait l'index par valeur) |
+6
View File
@@ -1056,6 +1056,12 @@ export function renderXlsxViewer(area, data) {
const dirtyCount = () => area.querySelectorAll("td.xlsx-dirty").length; const dirtyCount = () => area.querySelectorAll("td.xlsx-dirty").length;
const refreshSaveState = () => { saveBtn.disabled = dirtyCount() === 0; }; const refreshSaveState = () => { saveBtn.disabled = dirtyCount() === 0; };
// #153 A12 — the backend marks the last value Excel computed; the wording is
// translated here so the tooltip follows the UI language.
area.querySelectorAll(".xlsx-cached[data-cached-value]").forEach((el) => {
el.title = t("xlsx.cached_value_title");
});
// Editable cells: Enter blurs, Escape reverts, paste stays single-line. // Editable cells: Enter blurs, Escape reverts, paste stays single-line.
area.querySelectorAll(".xlsx-table td").forEach((td) => { area.querySelectorAll(".xlsx-table td").forEach((td) => {
td.contentEditable = "true"; td.contentEditable = "true";
+1
View File
@@ -1828,6 +1828,7 @@
"xlsx.lossy_confirm": "Save anyway? The following will be lost: {features}", "xlsx.lossy_confirm": "Save anyway? The following will be lost: {features}",
"xlsx.lossy_cancelled": "Save cancelled", "xlsx.lossy_cancelled": "Save cancelled",
"xlsx.formula_toggle_title": "Treat “=” and “@” as formulas (off by default)", "xlsx.formula_toggle_title": "Treat “=” and “@” as formulas (off by default)",
"xlsx.cached_value_title": "Last value calculated by Excel",
"xlsx.feature_cached_values": "cached values", "xlsx.feature_cached_values": "cached values",
"xlsx.feature_slicers": "slicers and timelines", "xlsx.feature_slicers": "slicers and timelines",
"xlsx.feature_form_controls": "form controls", "xlsx.feature_form_controls": "form controls",
+1
View File
@@ -1828,6 +1828,7 @@
"xlsx.lossy_confirm": "Enregistrer quand même ? Les éléments suivants seront perdus : {features}", "xlsx.lossy_confirm": "Enregistrer quand même ? Les éléments suivants seront perdus : {features}",
"xlsx.lossy_cancelled": "Sauvegarde annulée", "xlsx.lossy_cancelled": "Sauvegarde annulée",
"xlsx.formula_toggle_title": "Interpréter « = » et « @ » comme des formules (désactivé par défaut)", "xlsx.formula_toggle_title": "Interpréter « = » et « @ » comme des formules (désactivé par défaut)",
"xlsx.cached_value_title": "Dernière valeur calculée par Excel",
"xlsx.feature_cached_values": "valeurs calculées", "xlsx.feature_cached_values": "valeurs calculées",
"xlsx.feature_slicers": "segments et chronologies", "xlsx.feature_slicers": "segments et chronologies",
"xlsx.feature_form_controls": "contrôles de formulaire", "xlsx.feature_form_controls": "contrôles de formulaire",
+14
View File
@@ -10987,6 +10987,20 @@ body.desktop-mode .editor-container {
background: rgba(255, 196, 0, 0.18); background: rgba(255, 196, 0, 0.18);
} }
/* #153 A12 — last result Excel computed, shown under a formula cell.
Discreet by design: the formula is what the user edits, the cached value is
context (stale until Excel recalculates). */
.xlsx-cached {
display: block;
margin-top: 2px;
padding-left: 6px;
border-left: 2px solid var(--border, #d0d7de);
color: var(--text-muted);
font-size: 0.85em;
font-variant-numeric: tabular-nums;
white-space: nowrap;
}
/* #153 A1/A4 — lossy-save warning + formula toggle */ /* #153 A1/A4 — lossy-save warning + formula toggle */
.xlsx-warning { .xlsx-warning {
display: flex; display: flex;
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "obsigate", "name": "obsigate",
"version": "2.29.0", "version": "2.30.0",
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.", "description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
"main": "patch.js", "main": "patch.js",
"directories": { "directories": {
+25 -2
View File
@@ -11,7 +11,8 @@
* - the warning banner lists both features ; * - the warning banner lists both features ;
* - saving a cell on that workbook asks for confirmation (native dialog) and * - saving a cell on that workbook asks for confirmation (native dialog) and
* then succeeds (the client retries with `force: true`) ; * then succeeds (the client retries with `force: true`) ;
* - the f(x) toggle is off by default, so "=B1*3" is stored as text. * - the f(x) toggle is off by default, so "=B1*3" is stored as text ;
* - the value Excel last computed is shown under the formula (#153 A12).
* *
* The fixture is restored byte-for-byte in `afterAll` so a local run never * The fixture is restored byte-for-byte in `afterAll` so a local run never
* dirties the working copy. * dirties the working copy.
@@ -53,7 +54,7 @@ async function openFixture(page) {
await expect(page.locator('#content-area .xlsx-table')).toBeVisible({ timeout: 15000 }); await expect(page.locator('#content-area .xlsx-table')).toBeVisible({ timeout: 15000 });
} }
test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => { test.describe('Excel viewer — garde-fous d\'écriture et valeurs calculées (#153)', () => {
test.beforeAll(() => { test.beforeAll(() => {
if (existsSync(FIXTURE_PATH)) originalBytes = readFileSync(FIXTURE_PATH); if (existsSync(FIXTURE_PATH)) originalBytes = readFileSync(FIXTURE_PATH);
}); });
@@ -62,6 +63,28 @@ test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
if (originalBytes) writeFileSync(FIXTURE_PATH, originalBytes); if (originalBytes) writeFileSync(FIXTURE_PATH, originalBytes);
}); });
// Read-only assertions come FIRST, before the mutating tests: saving through
// the viewer rewrites the workbook and an openpyxl round-trip drops the cached
// formula results (BUG-085), so the shadow line only exists on a pristine
// fixture.
test('affiche la valeur calculée en cache sous la formule (#153 A12)', async ({ page }) => {
await login(page);
await openFixture(page);
// B2 is "=B1*2" and the package keeps its last result (<v>200</v>).
const formulaCell = page.locator('#content-area td[data-cell="B2"]');
await expect(formulaCell).toContainText('=B1*2');
const cached = formulaCell.locator('.xlsx-cached');
await expect(cached).toHaveCount(1);
await expect(cached).toHaveText('200');
// The tooltip is translated client-side, never hardcoded by the backend.
await expect(cached).toHaveAttribute('title', /Excel/);
// A plain value cell must not be duplicated with a shadow line.
await expect(page.locator('#content-area td[data-cell="B1"] .xlsx-cached')).toHaveCount(0);
});
test('affiche la bannière listant les éléments non préservés', async ({ page }) => { test('affiche la bannière listant les éléments non préservés', async ({ page }) => {
await login(page); await login(page);
await openFixture(page); await openFixture(page);
+194 -2
View File
@@ -106,16 +106,208 @@ class TestXlsxIndexing:
assert ".xlsx" in SUPPORTED_EXTENSIONS assert ".xlsx" in SUPPORTED_EXTENSIONS
def test_xlsx_indexed_metadata_only(self, test_vault_dir, xlsx_file): def test_xlsx_indexes_sheet_names_and_headers(self, test_vault_dir, xlsx_file):
"""#153 A5 — a workbook is searchable by its cell values."""
from backend.indexer import _index_single_file_sync from backend.indexer import _index_single_file_sync
info = _index_single_file_sync(VAULT, test_vault_dir, xlsx_file) info = _index_single_file_sync(VAULT, test_vault_dir, xlsx_file)
assert info is not None assert info is not None
assert info["extension"] == ".xlsx" assert info["extension"] == ".xlsx"
assert info["content"] == "" # binary: never read into TF-IDF # Sheet names + header rows reach TF-IDF (was metadata-only before A5).
assert "Budget" in info["content"]
assert "Poste" in info["content"]
assert info["content_preview"]
assert info["title"] # filename-derived title assert info["title"] # filename-derived title
class TestXlsxCachedValues:
"""#153 A12 — show what Excel last computed next to each formula."""
def test_cached_result_is_shown_beside_the_formula(self, client, lossy_xlsx):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(lossy_xlsx))[0]["html"]
# A2 is "=A1*3" with <v>9</v> in the fixture.
assert 'data-cell="A2"' in html
assert "=A1*3" in html
assert "xlsx-cached" in html
assert ">9<" in html # the cached result Excel computed
def test_no_shadow_when_no_formula_carries_a_result(self, client, xlsx_file):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(xlsx_file))[0]["html"]
assert "xlsx-cached" not in html # budget.xlsx has no <v> at all
def test_plain_cells_are_never_duplicated(self, client, lossy_xlsx):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(lossy_xlsx))[0]["html"]
# A1 is the literal 3: the two reads agree, so only one value shows.
assert 'data-cell="A1"' in html
assert html.count("xlsx-cached") == 1 # only the formula cell
class TestXlsxValueCoercion:
"""#153 A10 — a typed value comes back with the type Excel would infer."""
def _write(self, client, xlsx_file, ref, value):
return client.put(
f"/api/file/{VAULT}/xlsx/save",
params={"path": "budget.xlsx"},
json={"sheet": "Budget", "cells": {ref: value}, "force": True},
)
def test_number_and_bool_are_stored_as_typed(self, client, xlsx_file):
resp = self._write(
client, xlsx_file, "D1", "42"
)
assert resp.status_code == 200
resp = self._write(client, xlsx_file, "D2", "VRAI")
assert resp.status_code == 200
resp = self._write(client, xlsx_file, "D3", "12/03/2026")
assert resp.status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
ws = wb["Budget"]
assert ws["D1"].value == 42 and isinstance(ws["D1"].value, int)
assert ws["D2"].value is True
assert ws["D3"].value.year == 2026 and ws["D3"].value.month == 3
assert ws["D3"].value.day == 12 # FR day-first, not 3 December
wb.close()
def test_day_first_date_is_not_read_as_us(self, client, xlsx_file):
"""'01/02/2026' is 1 February in French, not 2 January."""
assert self._write(client, xlsx_file, "E1", "01/02/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
d = wb["Budget"]["E1"].value
wb.close()
assert (d.month, d.day) == (2, 1)
def test_ambiguous_text_is_left_alone(self, client, xlsx_file):
"""A version string or a partial date stays text, never a date."""
assert self._write(client, xlsx_file, "F1", "3.14.2").status_code == 200
assert self._write(client, xlsx_file, "F2", "Ref 12/34").status_code == 200
assert self._write(client, xlsx_file, "F3", "12/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
ws = wb["Budget"]
assert ws["F1"].value == "3.14.2"
assert ws["F2"].value == "Ref 12/34"
assert ws["F3"].value == "12/2026"
wb.close()
def test_plain_integer_text_becomes_a_number(self, client, xlsx_file):
"""Typing a bare number yields a number, as it did before #153 A10."""
assert self._write(client, xlsx_file, "H1", "75001").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
assert wb["Budget"]["H1"].value == 75001
wb.close()
def test_formula_looking_date_stays_text(self, client, xlsx_file):
"""A formula is never mistaken for a date (BUG-088 must not regress)."""
assert self._write(client, xlsx_file, "G1", "=12/03/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
c = wb["Budget"]["G1"]
wb.close()
assert c.data_type == "s"
assert c.value == "=12/03/2026"
class TestXlsxSearchable:
"""#153 A5 — a keyword living in a CELL must make the file findable."""
def test_search_finds_a_word_stored_in_a_cell(self, test_vault_dir):
"""A word typed in a CELL must make the workbook findable.
Deliberately hermetic: it drives the indexer and the inverted index
directly instead of going through the HTTP reload, because both are
process-wide singletons that other test modules mutate (some reload the
module outright), which would make this test order-dependent.
"""
from openpyxl import Workbook
import backend.indexer as indexer
import backend.search as search_mod
path = Path(test_vault_dir) / "fournisseurs.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "Contacts"
ws["A1"] = "Fournisseur"
ws["A2"] = "Menuiserie Beaulieu"
wb.save(path)
# Index exactly this vault, from scratch.
indexer.index[VAULT] = {
"files": [indexer._index_single_file_sync(VAULT, test_vault_dir, str(path))],
"tags": {},
"paths": [],
"config": {},
}
entries = [f for f in indexer.index[VAULT]["files"] if f]
assert entries, "le tableur n'a pas ete indexe"
assert "Menuiserie" in entries[0]["content"], (
f"contenu indexe : {entries[0]['content'][:80]!r}"
)
search_mod.init_inverted_index()
hits = [r["path"] for r in search_mod.search("Menuiserie", "all")]
assert "fournisseurs.xlsx" in hits
def test_indexable_text_is_capped(self, tmp_path):
"""A data dump must not flood the index."""
from openpyxl import Workbook
from backend.xlsx_reader import MAX_INDEX_CHARS, extract_indexable_text
path = tmp_path / "huge.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "Big"
for r in range(1, 400):
ws.cell(row=r, column=1, value=f"ligne {r} " + "x" * 60)
wb.save(path)
text = extract_indexable_text(path)
assert len(text) <= MAX_INDEX_CHARS
assert "Big" in text # the sheet name survives
def test_corrupt_workbook_indexes_as_empty_not_a_crash(self, tmp_path):
from backend.xlsx_reader import extract_indexable_text
path = tmp_path / "broken.xlsx"
path.write_bytes(b"not a zip at all")
assert extract_indexable_text(path) == ""
def test_blank_rows_are_skipped(self, tmp_path):
from openpyxl import Workbook
from backend.xlsx_reader import extract_indexable_text
path = tmp_path / "sparse.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "S"
ws["A1"] = "Alpha"
ws["A50"] = "Omega" # beyond the indexed prefix -> ignored on purpose
wb.save(path)
text = extract_indexable_text(path)
assert "Alpha" in text
assert "Omega" not in text
assert "\t\t" not in text # no run of empty columns
# ── Edit ────────────────────────────────────────────────────────────────── # ── Edit ──────────────────────────────────────────────────────────────────