feat: tableurs indexables, saisie typée et valeurs calculées #153

Les tableurs étaient invisibles à la recherche, chaque saisie devenait
du texte, et une formule n'affichait que sa formule.

- A5 : extract_indexable_text() indexe les noms de feuilles et les 20
  premières lignes (plafond 5 k caractères) pour le TF-IDF et la
  recherche sémantique. Un mot tapé dans une cellule rend le fichier
  trouvable ; un classeur chiffré s'indexe par son seul nom.
- A10 : la saisie est typée comme dans Excel — booléens
  (TRUE/FAUX/OUI/NON) et dates FR JJ/MM/AAAA en ordre jour-first, donc
  01/02/2026 est le 1er février. Une saisie ressemblant à une formule
  n'est jamais convertie (BUG-088 préservé).
- A12 : la valeur calculée par Excel s'affiche sous la formule, via une
  2e lecture data_only=True faite seulement si l'archive contient un
  <v>. Info-bulle traduite FR/EN, aucun texte d'interface côté backend.
- BUG-089 : un reindex manuel ne reconstruisait pas l'index inversé, et
  backend/search.py lisait l'index via un import par valeur — après un
  rechargement du module, la recherche écrivait dans un dict périmé.

Contre-preuves vérifiées pour A5, A10 et A12 (neutralisation de chaque
fonction → échec des tests concernés). Tests : 1402 pytest, 10 JSDOM,
4 E2E, suite E2E complète verte, ruff/mypy 0.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <[email protected]>
This commit is contained in:
2026-09-27 22:43:24 -04:00
parent 31d4616baf
commit 06f8e63d06
21 changed files with 585 additions and 61 deletions
+41 -1
View File
@@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/),
et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
> **En cours de développement** : les changements à venir sont listés dans la section
> [Unreleased](#unreleased). La dernière version livrée est **2.29.0**.
> [Unreleased](#unreleased). La dernière version livrée est **2.30.0**.
---
@@ -14,6 +14,46 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
---
## [2.30.0] — 2026-09-27
### Correction
- **BUG-089 — un reindex manuel ne reconstruisait pas l'index inversé.**
`reload_index()` / `reload_single_vault()` remplacent l'entrée de vault
en bloc, ce qui n'émet pas les notifications incrémentales : la
recherche TF-IDF continuait de servir un index périmé après un
reindex. Les deux fonctions appellent désormais `init_inverted_index()`.
Au passage, `backend/search.py` lisait l'index via
`from backend.indexer import index` — une liaison **par valeur** du
dict : un rechargement du module `backend.indexer` recréait le dict
côté indexer alors que la recherche écrivait dans l'ancien, et
l'index inversé n'indexait plus rien. Tous les accès passent par
`_indexer.index`. *Trouvé en écrivant le test de recherche d'A5 : il
passait isolément et échouait en suite complète selon l'ordre.*
### Ajouté
- **#153 A5 — les tableurs sont indexés par leur contenu.**
`extract_indexable_text()` extrait les noms de feuilles et les 20
premières lignes (plafond 5 000 caractères, 20 feuilles) pour le
TF-IDF et la recherche sémantique. Un mot tapé dans une cellule rend
désormais le classeur trouvable ; la lecture binaire pour l'affichage
est inchangée et un classeur chiffré/corrompu s'indexe par son seul
nom.
- **#153 A10 — la saisie est typée comme dans Excel.** Une valeur
`TRUE`/`FAUX`/`OUI`/`NON` devient un booléen, une date `JJ/MM/AAAA`
(avec `HH:MM` optionnel) devient une vraie date — et dans l'ordre
français : `01/02/2026` est le 1ᵉʳ février. Une saisie ressemblant à
une formule n'est jamais convertie.
- **#153 A12 — la valeur calculée s'affiche sous la formule.** Quand une
cellule porte encore le résultat de son dernier calcul Excel, celui-ci
s'affiche dans une ligne discrète sous la formule. La seconde lecture
`data_only=True` n'a lieu que si l'archive contient réellement une
valeur en cache, et toute erreur retombe sur l'affichage formules seul.
Info-bulle traduite FR/EN (`xlsx.cached_value_title`).
---
## [2.29.0] — 2026-09-27
### Correction
+3 -3
View File
@@ -4,7 +4,7 @@
**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.
[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -976,8 +976,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l
## 📝 Changelog
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.29.0).
Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.30.0).
---
*Projet : ObsiGate | Version : 2.29.0 | Dernière mise à jour : Septembre 2026*
*Projet : ObsiGate | Version : 2.30.0 | Dernière mise à jour : Septembre 2026*
+3 -3
View File
@@ -2,7 +2,7 @@
**Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface.
[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]()
[![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]()
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/)
[![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/)
@@ -1151,8 +1151,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE)
## 📝 Changelog
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.29.0).
See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.30.0).
---
*Project: ObsiGate | Version: 2.29.0 | Last updated: September 2026*
*Project: ObsiGate | Version: 2.30.0 | Last updated: September 2026*
+1 -1
View File
@@ -1 +1 @@
2.29.0
2.30.0
+39 -7
View File
@@ -351,6 +351,23 @@ def _decompress_excalidraw(compressed: str) -> dict[str, Any] | None:
return data
def extract_xlsx_indexable(file_path: Path) -> str:
"""Return searchable text for a workbook (#153 A5).
Lazy wrapper: ``openpyxl`` is only imported when a spreadsheet is actually
indexed, so a vault without workbooks never pays the import. Errors are
swallowed — a corrupt or encrypted file still gets indexed by name.
"""
try:
from backend.xlsx_reader import extract_indexable_text
except Exception: # pragma: no cover - openpyxl missing
return ""
try:
return extract_indexable_text(file_path)
except Exception: # pragma: no cover - defensive
return ""
def extract_excalidraw_indexable(raw: str) -> str:
"""Return indexable text content for a raw .excalidraw / .excalidraw.md file.
@@ -561,11 +578,12 @@ def _scan_vault(
title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = ""
elif ext == ".xlsx":
# #152 — binary workbook: metadata only, the viewer renders
# it (parity with _index_single_file_sync).
raw = ""
# #153 A5 — a workbook stays rendered by the viewer, but its
# cell values are now indexed as text so a spreadsheet is
# findable by its content (parity with _index_single_file_sync).
raw = extract_xlsx_indexable(fpath)
title = fpath.stem.replace("-", " ").replace("_", " ")
content_preview = ""
content_preview = raw[:200].strip()
else:
raw = fpath.read_text(encoding="utf-8", errors="replace")
title = fpath.stem.replace("-", " ").replace("_", " ")
@@ -807,6 +825,13 @@ async def reload_index() -> dict[str, Any]:
await build_index()
# BUG-040/#86: complete the deferred PDF + excalidraw extraction.
await enrich_pdf_texts()
# The inverted index is NOT updated by the hooks here: the rebuild above
# replaces whole vault entries, so the incremental notifications are not
# emitted for the files that only changed content. Without this, a manual
# reindex left TF-IDF search serving a stale index (BUG-089).
from backend.search import init_inverted_index
init_inverted_index()
stats = {}
for name, data in index.items():
stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])}
@@ -882,6 +907,13 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]:
# BUG-040/#86: complete the deferred PDF + excalidraw extraction.
await enrich_pdf_texts(vault_name)
# Same as reload_index: the vault entry was replaced wholesale, so rebuild
# the inverted index or TF-IDF search keeps serving stale postings
# (BUG-089).
from backend.search import init_inverted_index
init_inverted_index()
stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])}
logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags")
return stats
@@ -955,9 +987,9 @@ def _index_single_file_sync(vault_name: str, vault_path: str, file_path: str, va
raw = ""
content_preview = ""
elif ext == ".xlsx":
# #152 — binary workbook: metadata only (parity with _scan_vault).
raw = ""
content_preview = ""
# #153 A5 — index sheet names + header rows as text (see _scan_vault).
raw = extract_xlsx_indexable(fpath)
content_preview = raw[:200].strip()
else:
raw = fpath.read_text(encoding="utf-8", errors="replace")
content_preview = raw[:200].strip()
+10 -5
View File
@@ -12,7 +12,12 @@ from sortedcontainers import SortedList
from backend import indexer as _indexer
from backend import semantic_search as _semantic
from backend.indexer import index
# NOTE: the shared index is read through ``_indexer.index`` everywhere, never
# via ``from backend.indexer import index``. That import binds the dict object
# once, so a module reload of ``backend.indexer`` (tests, dev reload) rebinds
# the module-level name to a FRESH dict while this module keeps writing to the
# stale one — the inverted index then silently indexes nothing (BUG-089).
from backend.services.regex_safety import (
MAX_REGEX_MATCHES,
truncate_for_regex,
@@ -399,7 +404,7 @@ class InvertedIndex:
self.vault_docs = defaultdict(set)
self.tag_docs = defaultdict(set)
for vault_name, vault_data in index.items():
for vault_name, vault_data in _indexer.index.items():
for file_info in vault_data.get("files", []):
doc_key = f"{vault_name}::{file_info['path']}"
self.doc_count += 1
@@ -688,7 +693,7 @@ _indexer.set_index_change_hook(_on_index_change_hook)
def init_inverted_index():
"""Force initial inverted index build. Called after build_index completes on startup."""
if any(vdata.get("files") for vdata in index.values()):
if any(vdata.get("files") for vdata in _indexer.index.values()):
_inverted_index.rebuild()
logger.info("Inverted index initialized.")
@@ -784,7 +789,7 @@ def search(
else:
candidates = [
(vault_name, file_info)
for vault_name, vault_data in index.items()
for vault_name, vault_data in _indexer.index.items()
if vault_filter == "all" or vault_name == vault_filter
for file_info in vault_data["files"]
]
@@ -1613,7 +1618,7 @@ def get_all_tags(vault_filter: str | None = None) -> dict[str, int]:
Dict mapping tag names to their total occurrence count.
"""
merged: dict[str, int] = {}
for vault_name, vault_data in index.items():
for vault_name, vault_data in _indexer.index.items():
if vault_filter and vault_filter != "all" and vault_name != vault_filter:
continue
for tag, count in vault_data.get("tags", {}).items():
+50 -1
View File
@@ -19,6 +19,7 @@ import shutil
import threading
from collections.abc import Callable, Iterator
from contextlib import contextmanager
from datetime import date, datetime
from pathlib import Path
from typing import Any
@@ -237,6 +238,14 @@ _XLSX_FLOAT_RE = re.compile(r"^[+-]?(?:\d+\.\d*|\.\d+)$")
# the legacy Lotus-style trigger. "+"/"-" are left alone: they are numbers here.
_XLSX_FORMULA_RE = re.compile(r"^[=@]")
# #153 A10 — types recognised when a user types into a cell. Excel infers them
# too; storing everything as text would make a spreadsheet unusable (a boolean
# column stays a string, a date column sorts lexicographically).
_XLSX_TRUE_LITERALS = {"true", "vrai", "oui", "yes"}
_XLSX_FALSE_LITERALS = {"false", "faux", "non", "no"}
# Shape check before strptime: keeps the hot path free of format attempts.
_XLSX_DATE_RE = re.compile(r"^\d{1,2}[-/]\d{1,2}[-/]\d{4}(?:[ T]\d{1,2}:\d{2})?$")
# #153 A3 — per-file write lock. Two concurrent saves (two tabs, the AI agent
# and the viewer, a watcher restore) would otherwise read-modify-write on the
# same archive and the last writer silently wins. Kept deliberately small: the
@@ -270,7 +279,21 @@ def _xlsx_write_lock(key: str) -> Iterator[None]:
def _coerce_xlsx_value(value: Any) -> Any:
"""Turn the string sent by the cell editor back into a scalar."""
"""Turn the string sent by the cell editor back into a scalar (#153 A10).
The coercion is symmetric with :func:`backend.xlsx_reader._fmt`: a value
typed by the user comes back as a string, and Excel would have inferred a
type when typing the same thing. Recognised here:
* an empty cell -> ``None`` (clears it)
* ``1234`` / ``-1`` -> ``int``
* ``1.5`` / ``.5`` -> ``float``
* ``TRUE``/``FAUX`` (case-insensitive) -> ``bool``
* ``31/12/2026`` / ``31/12/2026 14:30`` -> ``date``/``datetime`` (FR)
Anything else stays text. A date-looking string typed with a leading
``=`` is a formula and never reaches here as a date.
"""
if not isinstance(value, str):
return value
text = value.strip()
@@ -280,9 +303,35 @@ def _coerce_xlsx_value(value: Any) -> Any:
return int(text)
if _XLSX_FLOAT_RE.match(text):
return float(text)
lowered = text.lower()
if lowered in _XLSX_TRUE_LITERALS:
return True
if lowered in _XLSX_FALSE_LITERALS:
return False
if not _XLSX_FORMULA_RE.match(text):
parsed = _parse_fr_datetime(text)
if parsed is not None:
return parsed
return value
def _parse_fr_datetime(text: str) -> date | datetime | None:
"""Parse a FR-localised date/datetime, or return ``None``.
Accepts ``JJ/MM/AAAA`` and ``JJ/MM/AAAA HH:MM`` (also ``JJ-MM-AAAA``).
``dayfirst`` is what makes ``01/02/2026`` the 1st of February rather than
the 2nd of January — the French convention.
"""
if not _XLSX_DATE_RE.match(text):
return None
for fmt in ("%d/%m/%Y %H:%M", "%d/%m/%Y", "%d-%m-%Y %H:%M", "%d-%m-%Y"):
try:
return datetime.strptime(text, fmt)
except ValueError:
continue
return None
def _write_cell(ws: Any, ref: str, value: Any, *, allow_formula: bool) -> None:
"""Assign one cell, forcing text when it looks like a formula.
+163 -13
View File
@@ -11,6 +11,7 @@ round-trip would drop (#153 A1) so the UI can warn before saving.
from __future__ import annotations
import html
import logging
import re
import zipfile
from datetime import date, datetime
@@ -20,6 +21,8 @@ from typing import Any
from openpyxl import load_workbook
from openpyxl.utils import get_column_letter
logger = logging.getLogger("obsigate.xlsx_reader")
# ponytail: hard caps bound the rendered grid (500 rows x 40 cols per sheet).
# Raise them, or paginate per sheet, if a real workbook needs more.
MAX_ROWS = 500
@@ -48,6 +51,14 @@ _CACHED_FORMULA_RE = re.compile(rb"<f[ >][^<]*</f>\s*<v>[^<]")
# Sheet XML scanned by the cached-formula probe (CPU guard, like MAX_REPLACE_FILE_BYTES).
_MAX_PROBE_BYTES = 8_000_000
# #153 A5 — ceiling on the text handed to the TF-IDF / semantic index. A workbook
# is a data dump, not prose: indexing every cell would flood the inverted index
# and bury the notes. Sheet names + the first rows are enough to make a
# spreadsheet findable by its headers.
MAX_INDEX_CHARS = 5_000
_INDEX_ROWS_PER_SHEET = 20
MAX_INDEX_SHEETS = 20
def _fmt(value: Any) -> str:
if value is None:
@@ -74,7 +85,29 @@ def _trim(grid: list[list[str]]) -> list[list[str]]:
return [row[:width] for row in grid]
def _table(grid: list[list[str]]) -> str:
def _cell_cached(cached: list[list[str]] | None, r: int, c: int) -> str:
"""Return the cached result for a 0-based cell, or ``""``.
The shadow grid is read positionally and may be narrower than the formula
grid (``_trim`` collapses the trailing empty columns of each grid
independently), so every lookup is bounds-checked rather than assumed.
"""
if not cached or r >= len(cached):
return ""
row = cached[r]
return row[c] if c < len(row) else ""
def _table(grid: list[list[str]], cached: list[list[str]] | None = None) -> str:
"""Render a grid as an HTML table.
``cached`` is the same grid read with ``data_only=True`` (#153 A12): where a
formula cell still carries its last computed result, it is shown as a
discreet second line (``<span class="xlsx-cached">``) so the user sees the
number Excel last calculated instead of only the formula text. The span
carries ``data-cached-value`` and is titled client-side from
``xlsx.cached_value_title`` — the backend never emits UI text.
"""
if not grid:
return "<p><em>Feuille vide</em></p>"
n_cols = max(len(row) for row in grid)
@@ -90,7 +123,24 @@ def _table(grid: list[list[str]]) -> str:
out.append(f'<tr><th class="xlsx-rownum">{r}</th>')
for c, val in enumerate(row, start=1):
ref = f"{get_column_letter(c)}{r}"
out.append(f'<td data-cell="{ref}">{html.escape(val)}</td>')
# The cached result only makes sense for a formula cell: on a plain
# value cell the two reads are identical and showing both would
# duplicate the text.
shadow = ""
if cached is not None and val.startswith("="):
# `c` is 1-based (A1 notation) and `r` too, while the grid is
# 0-based: translate both.
cval = _cell_cached(cached, r - 1, c - 1)
if cval and cval != val:
# The tooltip is translated client-side from
# `xlsx.cached_value_title`; never hardcode UI text here.
shadow = (
f'<span class="xlsx-cached" data-cached-value="1">'
f"{html.escape(cval)}</span>"
)
out.append(
f'<td data-cell="{ref}">{html.escape(val)}{shadow}</td>'
)
out.append("</tr>")
out.append("</tbody></table></div>")
return "".join(out)
@@ -142,18 +192,118 @@ def inspect_workbook(file_path: Path) -> list[str]:
def render_sheets(file_path: Path) -> list[dict[str, str]]:
"""Return ``[{"name": sheet_title, "html": table_html}, ...]``."""
"""Return ``[{"name": sheet_title, "html": table_html}, ...]``.
Reads the workbook twice: once with ``data_only=False`` for the formulas
(what the user must edit) and, when any formula carries a cached result
(#153 A12), once with ``data_only=True`` to show what Excel last computed.
The second pass is skipped entirely when the archive holds no cached value,
so the common case still costs a single load.
"""
wb = load_workbook(str(file_path), read_only=True, data_only=False)
try:
sheets = []
for ws in wb.worksheets:
grid = [
[_fmt(v) for v in row]
for row in ws.iter_rows(
min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True
)
]
sheets.append({"name": ws.title, "html": _table(_trim(grid))})
return sheets
formulas = [_sheet_grid(ws) for ws in wb.worksheets]
titles = [ws.title for ws in wb.worksheets]
finally:
wb.close()
cached: list[list[list[str]]] | None = None
if _has_cached_values(file_path):
cached = _read_cached_grids(file_path, titles)
sheets = []
for i, title in enumerate(titles):
grid = _trim(formulas[i])
# The shadow grid is NOT trimmed independently: _trim drops the
# trailing empty columns of each grid on its own width, which would
# shift every cached value left of its formula. Indexing it
# positionally against the untrimmed grid keeps the two aligned.
shadow = cached[i] if cached is not None and i < len(cached) else None
sheets.append({"name": title, "html": _table(grid, shadow)})
return sheets
def _sheet_grid(ws: Any) -> list[list[str]]:
"""Read one worksheet into a grid of formatted strings, bounded by the caps."""
return [
[_fmt(v) for v in row]
for row in ws.iter_rows(
min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True
)
]
def _has_cached_values(file_path: Path) -> bool:
"""True when the archive holds at least one ``<f>…</f><v>…</v>``."""
try:
with zipfile.ZipFile(file_path) as zf:
return _has_cached_formulas(zf)
except (OSError, zipfile.BadZipFile):
return False
def _read_cached_grids(
file_path: Path, titles: list[str]
) -> list[list[list[str]]] | None:
"""Read every sheet with ``data_only=True`` (what Excel last computed).
Best effort: returns ``None`` on any failure so the viewer falls back to the
formula-only rendering. A workbook Excel opens but openpyxl cannot re-read
must still display.
"""
try:
wb = load_workbook(str(file_path), read_only=True, data_only=True)
except Exception:
return None
try:
grids = [_sheet_grid(ws) for ws in wb.worksheets]
if [ws.title for ws in wb.worksheets] != titles:
return None
return grids
except Exception:
logger.debug("xlsx cached values unavailable", exc_info=True)
return None
finally:
wb.close()
def extract_indexable_text(file_path: Path) -> str:
"""Return searchable text for the TF-IDF / semantic index (#153 A5).
Sheet names plus the first :data:`_INDEX_ROWS_PER_SHEET` rows of each
sheet, capped at :data:`MAX_INDEX_CHARS`. Rows are tab-joined so a search
for a header matches the sheet it belongs to.
Never raises: a corrupt, encrypted or unsupported workbook yields ``""`` so
the file still gets indexed by name (same contract as :func:`inspect_workbook`).
"""
chunks: list[str] = []
budget = MAX_INDEX_CHARS
try:
wb = load_workbook(str(file_path), read_only=True, data_only=True)
except Exception:
# Encrypted (BadZipFile) or not a real workbook: name-only indexing.
return ""
try:
for ws in wb.worksheets[:MAX_INDEX_SHEETS]:
if budget <= 0:
break
# The sheet title alone is a strong signal ("Recettes", "Budget").
block = [ws.title]
for row in ws.iter_rows(
min_row=1, max_row=_INDEX_ROWS_PER_SHEET, max_col=MAX_COLS, values_only=True
):
cells = [_fmt(v) for v in row]
# Skip blank rows instead of emitting runs of tabs.
if not any(c.strip() for c in cells):
continue
block.append("\t".join(cells).rstrip())
text = "\n".join(block)
chunks.append(text[:budget])
budget -= len(text)
except Exception:
# Truncated but still useful: keep whatever was collected.
pass
finally:
wb.close()
return "\n".join(c for c in chunks if c).strip()
+1 -1
View File
@@ -2626,7 +2626,7 @@ dependencies = [
[[package]]
name = "obsigate-desktop"
version = "2.29.0"
version = "2.30.0"
dependencies = [
"chrono",
"env_logger",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "obsigate-desktop"
version = "2.29.0"
version = "2.30.0"
description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian"
authors = ["Bruno Charest"]
edition = "2021"
+1 -1
View File
@@ -1,7 +1,7 @@
{
"$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json",
"productName": "ObsiGate",
"version": "2.29.0",
"version": "2.30.0",
"identifier": "com.obsigate.desktop",
"build": {
"frontendDist": "../frontend",
+2
View File
@@ -195,6 +195,7 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| *BUG-086* | Édition d'un `.xlsx` : `wb.save()` écrit en place, un plantage laisse un classeur corrompu | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::edit_xlsx_cells` | Simuler un `OSError` pendant `Workbook.save` → le fichier d'origine est tronqué | Écriture atomique : `wb.save(<nom>.<pid>.tmp)` puis `os.replace()` ; `.tmp` supprimé sur échec ; le backup `.bak` reste inchangé. Test : `TestXlsxAtomicWrite::test_failed_save_keeps_the_original` (octets identiques après échec) + `test_no_tmp_left_after_a_successful_save` | #153 A2. Le fichier temporaire a un suffixe `.tmp` → ignoré par le watcher (`_is_relevant` ne retient que les extensions supportées). Vérifié : cf. BUG-085 |
| *BUG-087* | Édition d'un `.xlsx` concurrente (deux onglets, agent IA + viewer) : read-modify-write sans verrou, le dernier écrivain gagne silencieusement | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::_xlsx_write_lock` | Deux `PUT xlsx/save` simultanés sur le même fichier → une écriture est écrasée sans trace | Verrou par chemin (registre + garde, timeout 15 s) autour du cycle load → edit → `os.replace` ; attente dépassée → **409** `conflict`. L'endpoint est devenu `def` (sync) pour que l'attente s'exécute dans le threadpool et ne bloque pas la boucle d'événements. Test : `TestXlsxWriteLock` (2) | #153 A3. Verrou en mémoire, par processus : protège les cas d'un même serveur (le cas desktop/Tauri). Vérifié : cf. BUG-085 |
| *BUG-088* | Injection de formule dans un `.xlsx` : une saisie `=cmd\|'/c calc'!A1` est stockée comme formule et s'exécute à l'ouverture dans Excel (DDE) | 🟢 corrigé | P0 | tableur Excel / sécurité | IA | `backend/services/mutations.py::_write_cell`, `backend/routers/files_write.py`, `frontend/js/viewer.js::renderXlsxViewer` | `PUT /api/file/V/xlsx/save` avec `{"sheet": "S", "cells": {"A1": "=1+1"}}` → la cellule sort en `data_type == "f"` | `cell.data_type = "s"` après affectation : le texte est stocké comme chaîne, aucun `<f>` n'est écrit. Opt-in via `allow_formula: true` (endpoint) et le bouton `f(x)` de la visionneuse (session, jamais persisté). Test : `TestXlsxFormulaGuard` (4) + `xlsx-viewer.test.mjs` (toggle) | #153 A4. `+`/`-` ne sont pas neutralisés : ils sont déjà convertis en nombre par `_coerce_xlsx_value`. Le handler global `ServiceError` expose désormais `code` + `details` (le client en a besoin pour le 409), et `api()` (frontend) les propage sur l'Error. Vérifié : cf. BUG-085 |
| *BUG-089* | Un reindex manuel ne reconstruisait pas l'index inversé : la recherche TF-IDF continuait de servir un index périmé | 🟢 corrigé | P1 | ⚙️ backend / recherche | IA | `backend/indexer.py::reload_index`, `backend/indexer.py::reload_single_vault`, `backend/search.py` | Modifier le contenu d'un fichier, puis `GET /api/index/reload` → la recherche renvoie encore l'ancien contenu (ou rien pour un fichier nouveau) | `reload_index()` / `reload_single_vault()` appellent `init_inverted_index()` après le rebuild (le remplacement wholesale d'une entrée de vault n'émet pas les notifications incrémentales). En prime, `backend/search.py` lisait l'index via `from backend.indexer import index` (liaison **par valeur** du dict) : un `importlib.reload(backend.indexer)` recréait le dict côté indexer tandis que la recherche écrivait encore dans l'ancien — l'index inversé n'indexait alors plus rien. Tous les accès passent désormais par `_indexer.index`. Contre-preuve : `TestXlsxSearchable::test_search_finds_a_word_stored_in_a_cell` échoue sans le correctif | #153 A5. Trouvé en écrivant le test de recherche d'A5 : il passait isolément et échouait en suite complète selon l'ordre. Le reload incrémental par fichier (watcher, edition) n'est pas concerné : il passe par le hook `_on_index_change`. Vérifié : suite 1402 passed / 6 skipped, ruff/mypy 0 |
### TODOs techniques (améliorations / nouvelles tâches)
@@ -212,6 +213,7 @@ Avant de corriger quoi que ce soit, un agent IA doit :
| Date | ID(s) traité(s) | Action | Fichiers modifiés | Résumé | Statut après |
|---|---|---|---|---|---|
| 2026-09-28 | BUG-089 (#153 A5, A10, A12) | Correction | `backend/xlsx_reader.py`, `backend/indexer.py`, `backend/search.py`, `backend/services/mutations.py`, `frontend/js/viewer.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/test_xlsx_viewer.py` | **Les tableurs deviennent visibles ettypés** : (A5) `extract_indexable_text()` indexe noms de feuilles + 20 premières lignes (plafond 5 k caractères) dans le TF-IDF et la recherche sémantique — un mot tapé dans une cellule rend le fichier trouvable ; (A10) `_coerce_xlsx_value()` reconnaît désormais les booléens (`TRUE`/`FAUX`/`OUI`/`NON`) et les dates FR `JJ/MM/AAAA` (jour-first : `01/02/2026` = 1er février), symétrique avec l'affichage ; (A12) la valeur calculée en cache s'affiche sous la formule (`<span class="xlsx-cached">`, 2ᵉ lecture `data_only=True` uniquement si l'archive contient un `<v>`), info-bulle traduite via `xlsx.cached_value_title` FR/EN. (BUG-089) un reindex manuel reconstruisait mal l'index inversé et `backend/search.py` lisait l'index par valeur. Contre-preuves vérifiées pour A5, A10 et A12. Vérifié : `test_xlsx_viewer.py` 43 passed, suite 1402 passed / 6 skipped, ruff 0, mypy 0, i18n parity, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10 | 🟢 corrigé (en attente vérif utilisateur) |
| 2026-09-27 | BUG-085 → BUG-088 (#153 A1-A4) | Correction | `backend/xlsx_reader.py`, `backend/services/mutations.py`, `backend/routers/files_read.py`, `backend/routers/files_write.py`, `backend/schemas.py`, `backend/main.py`, `frontend/js/viewer.js`, `frontend/js/auth.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `frontend/sw.js`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-lossy.xlsx`, `.gitea/workflows/ci.yml` | **Garde-fous d'écriture des classeurs Excel** : (BUG-085) `inspect_workbook()` détecte ce qu'un round-trip openpyxl perd (valeurs calculées, slicers, contrôles, connexions, custom XML, signature) → la lecture expose `xlsx_lossy_features`, la visionneuse affiche une bannière et `PUT xlsx/save` refuse sans `force` (**409** `xlsx_lossy_content`, confirmation explicite puis reprise) ; (BUG-086) écriture atomique `.tmp` + `os.replace` ; (BUG-087) verrou par fichier (409 `conflict`, endpoint sync pour le threadpool) ; (BUG-088) une saisie `=`/`@` est stockée en texte (`data_type = "s"`), sauf opt-in `allow_formula` / bouton `f(x)`. Le handler `ServiceError` expose désormais `code` + `details` et `api()` les propage. Périmètre de perte revalidé empiriquement sur openpyxl 3.1.5 (graphiques, images et TCD sont préservés). Vérifié : `test_xlsx_viewer.py` 31 passed, suite 1390 passed / 6 skipped, ruff/mypy 0, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10, E2E 3/3 | 🟢 corrigé (en attente vérif utilisateur) |
| *(exemple)* 2026-06-15 | BUG-001 | Correction | `frontend/app.js` | Réécriture de `renderFile()` pour préserver le DOM dashboard | 🟢 corrigé (en attente vérif) |
| 2026-09-09 | BUG-001, BUG-002 | Correction | `backend/main.py`, `frontend/excalidraw-editor.html`, `tests/test_pdf_stream.py` | BUG-001: Content-Disposition RFC 5987 (nom PDF accentué ne casse plus l'en-tête → plus de 500). BUG-002: suppression alias esm.sh (408 jotai) + React 19 cohérent + prop `excalidrawAPI` → Loading masqué, save OK. Vérifié: 534 tests backend verts + E2E navigateur. | 🟢 corrigé (en attente vérif utilisateur) |
+6 -6
View File
@@ -1,6 +1,6 @@
# ObsiGate — Roadmap
> **Version :** 2.29.0 | **Dernière mise à jour :** 2026-09-27
> **Version :** 2.30.0 | **Dernière mise à jour :** 2026-09-27
> **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact
> vers les fonctionnalités livrées.
> - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)**
@@ -47,7 +47,7 @@
### 153. Visionneuse & édition XLSX — complétude (fidélité, recherche, IA, UX, formats)
- **Effort :** 8-13 jours (P0 ✅ 2-3 j · P1 : 4-6 j · P2 : 2-4 j) | **Impact :** 🟡
- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), reste P1 puis P2
- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17
- **Analyse, risques et critères d'acceptation :** [features/xlsx-viewer.md](./features/xlsx-viewer.md)
- **Description :** #152 (visionneuse XLSX, 2.27.0) lit et édite correctement la **grille de
valeurs** d'un `.xlsx`, mais l'ensemble supporté est étroit : valeurs seulement (ni structure,
@@ -70,15 +70,15 @@
- [x] **A2** Écriture atomique (`wb.save(.tmp)` + `os.replace()`, backup inchangé) — BUG-086
- [x] **A3** Verrou par fichier autour du read-modify-write (timeout 15 s + **409** `conflict`) — BUG-087
- [x] **A4** Neutralisation de l'injection de formule (`=`/`@` stockés en texte, opt-in `allow_formula` + bouton `f(x)`) — BUG-088
- **P1 — recherche, IA, UX (🟡, 4-6 j) — ⚪ à faire**
- [ ] **A5** Indexation du contenu des feuilles (TF-IDF + sémantique, plafond ~5 k caractères)
- **P1 — recherche, IA, UX (🟡, 4-6 j) — 🔵 en cours**
- [x] **A5** Indexation du contenu des feuilles (noms de feuilles + 20 premières lignes, plafond 5 k caractères) — les mots tapés dans une cellule rendent le fichier trouvable ; au passage **BUG-089** (reindex manuel ne reconstruisait pas l'index inversé)
- [ ] **A6** Outils IA `update_xlsx_cells` / `append_xlsx_rows` / `xlsx_to_markdown` / `list_xlsx_sheets`
- [ ] **A7** Navigation clavier + barre de formule + nom de cellule (Tab/Entrée/flèches, `Maj+Entrée`, copie de plage)
- [ ] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse)
- [ ] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel)
- [ ] **A10** Types & formats de saisie (nombre/texte, booléens, dates localisées FR)
- [x] **A10** Types & formats de saisie (nombre/texte, booléens `TRUE`/`FAUX`, dates FR `JJ/MM/AAAA` jour-first)
- [ ] **A11** Tests frontend (`tests/frontend/xlsx-viewer.test.mjs`) + E2E (`tests/e2e/xlsx-viewer.spec.js`) au CI
- [ ] **A12** Affichage de la valeur calculée en cache (lecture `data_only=True`, mention FR/EN)
- [x] **A12** Valeur calculée affichée sous la formule (2ᵉ lecture `data_only=True` seulement si l'archive contient un `<v>`, info-bulle FR/EN)
- **P2 — étendu (🟢, 2-4 j) — ⚪ à faire**
- [ ] **A13** Tri / filtre / recherche dans la feuille + export CSV de la sélection
- [ ] **A14** CRUD de feuilles, lignes et colonnes (renommer, insérer, supprimer, dupliquer)
+22 -13
View File
@@ -3,7 +3,7 @@
> **Item de roadmap :** [#153 — Visionneuse & édition XLSX — complétude](../ROADMAP.md)
> **Origine :** #152 (visionneuse XLSX, livrée en 2.27.0 — voir
> [archive/COMPLETED_v1-v2.md](../archive/COMPLETED_v1-v2.md))
> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), P1/P2 restants
> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17
> **Effort estimé :** 8-13 jours au total (P0 ✅ 2-3 j · P1 4-6 j · P2 2-4 j)
> **Règle de maintenance :** la Roadmap porte les cases à cocher (suivi), cette fiche porte
> l'analyse, les risques et les critères d'acceptation. **Ne pas dupliquer le détail.**
@@ -134,13 +134,14 @@ couverture) · effort en jours-homme de développement + tests.
nombres. Au passage : le handler `ServiceError` expose `code` + `details` et `api()` les
propage sur l'Error. *Vérifié :* `TestXlsxFormulaGuard` (4) + test du toggle côté UI.
### P1 — Recherche, IA, UX (4-6 j)
### P1 — Recherche, IA, UX (4-6 j) — 🟢 A5, A10, A12 livrés le 2026-09-28
- [ ] **A5 — Indexation du contenu des feuilles.** Extraire un texte (noms de feuilles +
en-têtes + N premières lignes, plafond ~5 k caractères) pour le TF-IDF et la recherche
sémantique, tout en gardant la lecture binaire pour l'affichage ; `content_preview`
renseigné ; exclusion si le classeur est chiffré/corrompu. *Critère :* une cellule contenant
un mot-clé rend le fichier trouvable ; `test_xlsx_indexing` étendu.
- [x] **A5 — Indexation du contenu des feuilles.** `extract_indexable_text()` (noms de feuilles +
20 premières lignes, `MAX_INDEX_CHARS = 5 000`, 20 feuilles max) alimente le TF-IDF et la
recherche sémantique ; la lecture binaire reste inchangée pour l'affichage. Un classeur
chiffré/corrompu s'indexe par son seul nom (jamais d'exception). Au passage : **BUG-089**,
un reindex manuel ne reconstruisait pas l'index inversé. *Vérifié :* `TestXlsxSearchable` (4)
+ `TestXlsxIndexing`, **contre-preuve** (neutraliser l'extraction → 3 tests échouent).
- [ ] **A6 — Outils IA sur classeur.** `update_xlsx_cells` (enveloppe du service existant),
`append_xlsx_rows`, `xlsx_to_markdown` (contexte LLM, plafonné), `list_xlsx_sheets` — risque
WRITE + confirmation pour les mutations, libellés i18n dans `backend/tools/labels.py`,
@@ -155,15 +156,22 @@ couverture) · effort en jours-homme de développement + tests.
`GET /api/file/{vault}/xlsx/sheet?sheet=N&offset=&limit=` (`response_model` +
`backend/openapi_docs.py`), rendu à la demande avec défilement virtuel, bouton « charger
tout ».
- [ ] **A10 — Types et formats de saisie.** Coercion symétrique à l'écriture/à l'affichage
(nombre vs texte, booléens `TRUE`/`FALSE`, dates localisées FR — le TODO existe déjà dans
`_coerce_xlsx_value`) ; affichage du type d'origine dans l'info-bulle de cellule.
- [x] **A10 — Types et formats de saisie.** `_coerce_xlsx_value()` reconnait les booléens
(`true`/`vrai`/`oui`/`yes` et leurs négatifs) et les dates FR `JJ/MM/AAAA` (+ `HH:MM`),
jour-first comme Excel en locale française : `01/02/2026` = 1ᵉʳ février. Une saisie
ressemblant à une formule n'est jamais convertie (BUG-088 préservé) ; un code postal
numérique ou une version restent ce qu'ils sont. *Vérifié :* `TestXlsxValueCoercion` (5),
**contre-preuve** (neutraliser la coercion → 2 tests échouent).
- [ ] **A11 — Tests frontend + E2E.** `tests/frontend/xlsx-viewer.test.mjs` (dirty, Échap,
collage, 1 PUT par feuille, bouton désactivé) et `tests/e2e/xlsx-viewer.spec.js`
(ouverture, onglets, édition, sauvegarde, rechargement) ; intégration au CI.
- [ ] **A12 — Valeurs calculées.** Afficher la valeur en cache (2ᵉ ligne discrète) quand elle
existe, via une lecture `data_only=True` de la même page d'onglets ; mention FR/EN
« valeur recalculée par Excel ».
- [x] **A12 — Valeurs calculées.** La valeur en cache s'affiche sous la formule dans un
`<span class="xlsx-cached">`. La 2ᵉ lecture `data_only=True` n'a lieu que si l'archive
contient réellement un `<f>…</f><v>…</v>` (sonde déjà présente pour A1) : le cas courant
reste à un seul chargement, et toute erreur retombe sur l'affichage formules seul.
L'info-bulle est traduite côté client (`xlsx.cached_value_title` FR/EN) — aucun texte
d'interface n'est émis par le backend. *Vérifié :* `TestXlsxCachedValues` (3),
**contre-preuve** (neutraliser la 2ᵉ lecture → 2 tests échouent).
### P2 — Étendu (2-4 j)
@@ -199,3 +207,4 @@ couverture) · effort en jours-homme de développement + tests.
| 2026-09-27 | Audit complet → création de #153 : limites, risques R1-R5, backlog A1-A17 |
| 2026-09-27 | Périmètre de perte **remesuré** sur openpyxl 3.1.5 : graphiques / images / TCD sont préservés, seules les valeurs en cache et quelques parties exotiques sont perdues |
| 2026-09-27 | **P0 livré** (BUG-085 → BUG-088) : `xlsx_lossy_features` + 409 `xlsx_lossy_content`, écriture atomique, verrou par fichier, formules stockées en texte par défaut |
| 2026-09-28 | **A5 + A10 + A12 livrés** : le contenu des cellules est indexé (recherche), la saisie est typée (booléens, dates FR), la valeur calculée s'affiche sous la formule. **BUG-089** corrigé au passage (reindex manuel ≠ reconstruction de l'index inversé ; `backend/search.py` lisait l'index par valeur) |
+6
View File
@@ -1056,6 +1056,12 @@ export function renderXlsxViewer(area, data) {
const dirtyCount = () => area.querySelectorAll("td.xlsx-dirty").length;
const refreshSaveState = () => { saveBtn.disabled = dirtyCount() === 0; };
// #153 A12 — the backend marks the last value Excel computed; the wording is
// translated here so the tooltip follows the UI language.
area.querySelectorAll(".xlsx-cached[data-cached-value]").forEach((el) => {
el.title = t("xlsx.cached_value_title");
});
// Editable cells: Enter blurs, Escape reverts, paste stays single-line.
area.querySelectorAll(".xlsx-table td").forEach((td) => {
td.contentEditable = "true";
+1
View File
@@ -1828,6 +1828,7 @@
"xlsx.lossy_confirm": "Save anyway? The following will be lost: {features}",
"xlsx.lossy_cancelled": "Save cancelled",
"xlsx.formula_toggle_title": "Treat “=” and “@” as formulas (off by default)",
"xlsx.cached_value_title": "Last value calculated by Excel",
"xlsx.feature_cached_values": "cached values",
"xlsx.feature_slicers": "slicers and timelines",
"xlsx.feature_form_controls": "form controls",
+1
View File
@@ -1828,6 +1828,7 @@
"xlsx.lossy_confirm": "Enregistrer quand même ? Les éléments suivants seront perdus : {features}",
"xlsx.lossy_cancelled": "Sauvegarde annulée",
"xlsx.formula_toggle_title": "Interpréter « = » et « @ » comme des formules (désactivé par défaut)",
"xlsx.cached_value_title": "Dernière valeur calculée par Excel",
"xlsx.feature_cached_values": "valeurs calculées",
"xlsx.feature_slicers": "segments et chronologies",
"xlsx.feature_form_controls": "contrôles de formulaire",
+14
View File
@@ -10987,6 +10987,20 @@ body.desktop-mode .editor-container {
background: rgba(255, 196, 0, 0.18);
}
/* #153 A12 — last result Excel computed, shown under a formula cell.
Discreet by design: the formula is what the user edits, the cached value is
context (stale until Excel recalculates). */
.xlsx-cached {
display: block;
margin-top: 2px;
padding-left: 6px;
border-left: 2px solid var(--border, #d0d7de);
color: var(--text-muted);
font-size: 0.85em;
font-variant-numeric: tabular-nums;
white-space: nowrap;
}
/* #153 A1/A4 — lossy-save warning + formula toggle */
.xlsx-warning {
display: flex;
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "obsigate",
"version": "2.29.0",
"version": "2.30.0",
"description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.",
"main": "patch.js",
"directories": {
+25 -2
View File
@@ -11,7 +11,8 @@
* - the warning banner lists both features ;
* - saving a cell on that workbook asks for confirmation (native dialog) and
* then succeeds (the client retries with `force: true`) ;
* - the f(x) toggle is off by default, so "=B1*3" is stored as text.
* - the f(x) toggle is off by default, so "=B1*3" is stored as text ;
* - the value Excel last computed is shown under the formula (#153 A12).
*
* The fixture is restored byte-for-byte in `afterAll` so a local run never
* dirties the working copy.
@@ -53,7 +54,7 @@ async function openFixture(page) {
await expect(page.locator('#content-area .xlsx-table')).toBeVisible({ timeout: 15000 });
}
test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
test.describe('Excel viewer — garde-fous d\'écriture et valeurs calculées (#153)', () => {
test.beforeAll(() => {
if (existsSync(FIXTURE_PATH)) originalBytes = readFileSync(FIXTURE_PATH);
});
@@ -62,6 +63,28 @@ test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => {
if (originalBytes) writeFileSync(FIXTURE_PATH, originalBytes);
});
// Read-only assertions come FIRST, before the mutating tests: saving through
// the viewer rewrites the workbook and an openpyxl round-trip drops the cached
// formula results (BUG-085), so the shadow line only exists on a pristine
// fixture.
test('affiche la valeur calculée en cache sous la formule (#153 A12)', async ({ page }) => {
await login(page);
await openFixture(page);
// B2 is "=B1*2" and the package keeps its last result (<v>200</v>).
const formulaCell = page.locator('#content-area td[data-cell="B2"]');
await expect(formulaCell).toContainText('=B1*2');
const cached = formulaCell.locator('.xlsx-cached');
await expect(cached).toHaveCount(1);
await expect(cached).toHaveText('200');
// The tooltip is translated client-side, never hardcoded by the backend.
await expect(cached).toHaveAttribute('title', /Excel/);
// A plain value cell must not be duplicated with a shadow line.
await expect(page.locator('#content-area td[data-cell="B1"] .xlsx-cached')).toHaveCount(0);
});
test('affiche la bannière listant les éléments non préservés', async ({ page }) => {
await login(page);
await openFixture(page);
+194 -2
View File
@@ -106,16 +106,208 @@ class TestXlsxIndexing:
assert ".xlsx" in SUPPORTED_EXTENSIONS
def test_xlsx_indexed_metadata_only(self, test_vault_dir, xlsx_file):
def test_xlsx_indexes_sheet_names_and_headers(self, test_vault_dir, xlsx_file):
"""#153 A5 — a workbook is searchable by its cell values."""
from backend.indexer import _index_single_file_sync
info = _index_single_file_sync(VAULT, test_vault_dir, xlsx_file)
assert info is not None
assert info["extension"] == ".xlsx"
assert info["content"] == "" # binary: never read into TF-IDF
# Sheet names + header rows reach TF-IDF (was metadata-only before A5).
assert "Budget" in info["content"]
assert "Poste" in info["content"]
assert info["content_preview"]
assert info["title"] # filename-derived title
class TestXlsxCachedValues:
"""#153 A12 — show what Excel last computed next to each formula."""
def test_cached_result_is_shown_beside_the_formula(self, client, lossy_xlsx):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(lossy_xlsx))[0]["html"]
# A2 is "=A1*3" with <v>9</v> in the fixture.
assert 'data-cell="A2"' in html
assert "=A1*3" in html
assert "xlsx-cached" in html
assert ">9<" in html # the cached result Excel computed
def test_no_shadow_when_no_formula_carries_a_result(self, client, xlsx_file):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(xlsx_file))[0]["html"]
assert "xlsx-cached" not in html # budget.xlsx has no <v> at all
def test_plain_cells_are_never_duplicated(self, client, lossy_xlsx):
from backend.xlsx_reader import render_sheets
html = render_sheets(Path(lossy_xlsx))[0]["html"]
# A1 is the literal 3: the two reads agree, so only one value shows.
assert 'data-cell="A1"' in html
assert html.count("xlsx-cached") == 1 # only the formula cell
class TestXlsxValueCoercion:
"""#153 A10 — a typed value comes back with the type Excel would infer."""
def _write(self, client, xlsx_file, ref, value):
return client.put(
f"/api/file/{VAULT}/xlsx/save",
params={"path": "budget.xlsx"},
json={"sheet": "Budget", "cells": {ref: value}, "force": True},
)
def test_number_and_bool_are_stored_as_typed(self, client, xlsx_file):
resp = self._write(
client, xlsx_file, "D1", "42"
)
assert resp.status_code == 200
resp = self._write(client, xlsx_file, "D2", "VRAI")
assert resp.status_code == 200
resp = self._write(client, xlsx_file, "D3", "12/03/2026")
assert resp.status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
ws = wb["Budget"]
assert ws["D1"].value == 42 and isinstance(ws["D1"].value, int)
assert ws["D2"].value is True
assert ws["D3"].value.year == 2026 and ws["D3"].value.month == 3
assert ws["D3"].value.day == 12 # FR day-first, not 3 December
wb.close()
def test_day_first_date_is_not_read_as_us(self, client, xlsx_file):
"""'01/02/2026' is 1 February in French, not 2 January."""
assert self._write(client, xlsx_file, "E1", "01/02/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
d = wb["Budget"]["E1"].value
wb.close()
assert (d.month, d.day) == (2, 1)
def test_ambiguous_text_is_left_alone(self, client, xlsx_file):
"""A version string or a partial date stays text, never a date."""
assert self._write(client, xlsx_file, "F1", "3.14.2").status_code == 200
assert self._write(client, xlsx_file, "F2", "Ref 12/34").status_code == 200
assert self._write(client, xlsx_file, "F3", "12/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
ws = wb["Budget"]
assert ws["F1"].value == "3.14.2"
assert ws["F2"].value == "Ref 12/34"
assert ws["F3"].value == "12/2026"
wb.close()
def test_plain_integer_text_becomes_a_number(self, client, xlsx_file):
"""Typing a bare number yields a number, as it did before #153 A10."""
assert self._write(client, xlsx_file, "H1", "75001").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
assert wb["Budget"]["H1"].value == 75001
wb.close()
def test_formula_looking_date_stays_text(self, client, xlsx_file):
"""A formula is never mistaken for a date (BUG-088 must not regress)."""
assert self._write(client, xlsx_file, "G1", "=12/03/2026").status_code == 200
from openpyxl import load_workbook
wb = load_workbook(xlsx_file)
c = wb["Budget"]["G1"]
wb.close()
assert c.data_type == "s"
assert c.value == "=12/03/2026"
class TestXlsxSearchable:
"""#153 A5 — a keyword living in a CELL must make the file findable."""
def test_search_finds_a_word_stored_in_a_cell(self, test_vault_dir):
"""A word typed in a CELL must make the workbook findable.
Deliberately hermetic: it drives the indexer and the inverted index
directly instead of going through the HTTP reload, because both are
process-wide singletons that other test modules mutate (some reload the
module outright), which would make this test order-dependent.
"""
from openpyxl import Workbook
import backend.indexer as indexer
import backend.search as search_mod
path = Path(test_vault_dir) / "fournisseurs.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "Contacts"
ws["A1"] = "Fournisseur"
ws["A2"] = "Menuiserie Beaulieu"
wb.save(path)
# Index exactly this vault, from scratch.
indexer.index[VAULT] = {
"files": [indexer._index_single_file_sync(VAULT, test_vault_dir, str(path))],
"tags": {},
"paths": [],
"config": {},
}
entries = [f for f in indexer.index[VAULT]["files"] if f]
assert entries, "le tableur n'a pas ete indexe"
assert "Menuiserie" in entries[0]["content"], (
f"contenu indexe : {entries[0]['content'][:80]!r}"
)
search_mod.init_inverted_index()
hits = [r["path"] for r in search_mod.search("Menuiserie", "all")]
assert "fournisseurs.xlsx" in hits
def test_indexable_text_is_capped(self, tmp_path):
"""A data dump must not flood the index."""
from openpyxl import Workbook
from backend.xlsx_reader import MAX_INDEX_CHARS, extract_indexable_text
path = tmp_path / "huge.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "Big"
for r in range(1, 400):
ws.cell(row=r, column=1, value=f"ligne {r} " + "x" * 60)
wb.save(path)
text = extract_indexable_text(path)
assert len(text) <= MAX_INDEX_CHARS
assert "Big" in text # the sheet name survives
def test_corrupt_workbook_indexes_as_empty_not_a_crash(self, tmp_path):
from backend.xlsx_reader import extract_indexable_text
path = tmp_path / "broken.xlsx"
path.write_bytes(b"not a zip at all")
assert extract_indexable_text(path) == ""
def test_blank_rows_are_skipped(self, tmp_path):
from openpyxl import Workbook
from backend.xlsx_reader import extract_indexable_text
path = tmp_path / "sparse.xlsx"
wb = Workbook()
ws = wb.active
ws.title = "S"
ws["A1"] = "Alpha"
ws["A50"] = "Omega" # beyond the indexed prefix -> ignored on purpose
wb.save(path)
text = extract_indexable_text(path)
assert "Alpha" in text
assert "Omega" not in text
assert "\t\t" not in text # no run of empty columns
# ── Edit ──────────────────────────────────────────────────────────────────