diff --git a/CHANGELOG.md b/CHANGELOG.md index 7d89465..0f72a28 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,7 +6,7 @@ Format basé sur [Keep a Changelog](https://keepachangelog.com/fr/1.1.0/), et [Semantic Versioning](https://semver.org/spec/v2.0.0.html). > **En cours de développement** : les changements à venir sont listés dans la section -> [Unreleased](#unreleased). La dernière version livrée est **2.29.0**. +> [Unreleased](#unreleased). La dernière version livrée est **2.30.0**. --- @@ -14,6 +14,46 @@ et [Semantic Versioning](https://semver.org/spec/v2.0.0.html). --- +## [2.30.0] — 2026-09-27 + +### Correction + +- **BUG-089 — un reindex manuel ne reconstruisait pas l'index inversé.** + `reload_index()` / `reload_single_vault()` remplacent l'entrée de vault + en bloc, ce qui n'émet pas les notifications incrémentales : la + recherche TF-IDF continuait de servir un index périmé après un + reindex. Les deux fonctions appellent désormais `init_inverted_index()`. + Au passage, `backend/search.py` lisait l'index via + `from backend.indexer import index` — une liaison **par valeur** du + dict : un rechargement du module `backend.indexer` recréait le dict + côté indexer alors que la recherche écrivait dans l'ancien, et + l'index inversé n'indexait plus rien. Tous les accès passent par + `_indexer.index`. *Trouvé en écrivant le test de recherche d'A5 : il + passait isolément et échouait en suite complète selon l'ordre.* + +### Ajouté + +- **#153 A5 — les tableurs sont indexés par leur contenu.** + `extract_indexable_text()` extrait les noms de feuilles et les 20 + premières lignes (plafond 5 000 caractères, 20 feuilles) pour le + TF-IDF et la recherche sémantique. Un mot tapé dans une cellule rend + désormais le classeur trouvable ; la lecture binaire pour l'affichage + est inchangée et un classeur chiffré/corrompu s'indexe par son seul + nom. +- **#153 A10 — la saisie est typée comme dans Excel.** Une valeur + `TRUE`/`FAUX`/`OUI`/`NON` devient un booléen, une date `JJ/MM/AAAA` + (avec `HH:MM` optionnel) devient une vraie date — et dans l'ordre + français : `01/02/2026` est le 1ᵉʳ février. Une saisie ressemblant à + une formule n'est jamais convertie. +- **#153 A12 — la valeur calculée s'affiche sous la formule.** Quand une + cellule porte encore le résultat de son dernier calcul Excel, celui-ci + s'affiche dans une ligne discrète sous la formule. La seconde lecture + `data_only=True` n'a lieu que si l'archive contient réellement une + valeur en cache, et toute erreur retombe sur l'affichage formules seul. + Info-bulle traduite FR/EN (`xlsx.cached_value_title`). + +--- + ## [2.29.0] — 2026-09-27 ### Correction diff --git a/README.fr.md b/README.fr.md index 3357061..8588467 100644 --- a/README.fr.md +++ b/README.fr.md @@ -4,7 +4,7 @@ **Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive. -[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]() +[![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]() [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/) [![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/) @@ -976,8 +976,8 @@ Ce projet est sous licence **MIT** — voir le fichier [LICENSE](LICENSE) pour l ## 📝 Changelog -Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.29.0). +Consultez le [CHANGELOG.md](./CHANGELOG.md) pour l'historique complet de toutes les versions (v1.0.0 → v2.30.0). --- -*Projet : ObsiGate | Version : 2.29.0 | Dernière mise à jour : Septembre 2026* +*Projet : ObsiGate | Version : 2.30.0 | Dernière mise à jour : Septembre 2026* diff --git a/README.md b/README.md index 8ba4248..544a355 100644 --- a/README.md +++ b/README.md @@ -2,7 +2,7 @@ **Ultra-light web gateway for your Obsidian vaults** — Access, browse, and search all your Obsidian notes from any device via a modern, responsive web interface. -[![Version](https://img.shields.io/badge/Version-2.29.0-blue.svg)]() +[![Version](https://img.shields.io/badge/Version-2.30.0-blue.svg)]() [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![Docker](https://img.shields.io/badge/Docker-Ready-blue.svg)](https://www.docker.com/) [![Python](https://img.shields.io/badge/Python-3.11+-green.svg)](https://www.python.org/) @@ -1151,8 +1151,8 @@ This project is licensed under the **MIT License** - see the [LICENSE](LICENSE) ## 📝 Changelog -See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.29.0). +See [CHANGELOG.md](./CHANGELOG.md) for the complete version history (v1.0.0 → v2.30.0). --- -*Project: ObsiGate | Version: 2.29.0 | Last updated: September 2026* +*Project: ObsiGate | Version: 2.30.0 | Last updated: September 2026* diff --git a/VERSION b/VERSION index f013568..6a69003 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -2.29.0 +2.30.0 diff --git a/backend/indexer.py b/backend/indexer.py index 6371229..5c7e2be 100644 --- a/backend/indexer.py +++ b/backend/indexer.py @@ -351,6 +351,23 @@ def _decompress_excalidraw(compressed: str) -> dict[str, Any] | None: return data +def extract_xlsx_indexable(file_path: Path) -> str: + """Return searchable text for a workbook (#153 A5). + + Lazy wrapper: ``openpyxl`` is only imported when a spreadsheet is actually + indexed, so a vault without workbooks never pays the import. Errors are + swallowed — a corrupt or encrypted file still gets indexed by name. + """ + try: + from backend.xlsx_reader import extract_indexable_text + except Exception: # pragma: no cover - openpyxl missing + return "" + try: + return extract_indexable_text(file_path) + except Exception: # pragma: no cover - defensive + return "" + + def extract_excalidraw_indexable(raw: str) -> str: """Return indexable text content for a raw .excalidraw / .excalidraw.md file. @@ -561,11 +578,12 @@ def _scan_vault( title = fpath.stem.replace("-", " ").replace("_", " ") content_preview = "" elif ext == ".xlsx": - # #152 — binary workbook: metadata only, the viewer renders - # it (parity with _index_single_file_sync). - raw = "" + # #153 A5 — a workbook stays rendered by the viewer, but its + # cell values are now indexed as text so a spreadsheet is + # findable by its content (parity with _index_single_file_sync). + raw = extract_xlsx_indexable(fpath) title = fpath.stem.replace("-", " ").replace("_", " ") - content_preview = "" + content_preview = raw[:200].strip() else: raw = fpath.read_text(encoding="utf-8", errors="replace") title = fpath.stem.replace("-", " ").replace("_", " ") @@ -807,6 +825,13 @@ async def reload_index() -> dict[str, Any]: await build_index() # BUG-040/#86: complete the deferred PDF + excalidraw extraction. await enrich_pdf_texts() + # The inverted index is NOT updated by the hooks here: the rebuild above + # replaces whole vault entries, so the incremental notifications are not + # emitted for the files that only changed content. Without this, a manual + # reindex left TF-IDF search serving a stale index (BUG-089). + from backend.search import init_inverted_index + + init_inverted_index() stats = {} for name, data in index.items(): stats[name] = {"file_count": len(data["files"]), "tag_count": len(data["tags"])} @@ -882,6 +907,13 @@ async def reload_single_vault(vault_name: str) -> dict[str, Any]: # BUG-040/#86: complete the deferred PDF + excalidraw extraction. await enrich_pdf_texts(vault_name) + # Same as reload_index: the vault entry was replaced wholesale, so rebuild + # the inverted index or TF-IDF search keeps serving stale postings + # (BUG-089). + from backend.search import init_inverted_index + + init_inverted_index() + stats = {"file_count": len(vault_data["files"]), "tag_count": len(vault_data["tags"])} logger.info(f"Vault '{vault_name}' reindexed: {stats['file_count']} files, {stats['tag_count']} tags") return stats @@ -955,9 +987,9 @@ def _index_single_file_sync(vault_name: str, vault_path: str, file_path: str, va raw = "" content_preview = "" elif ext == ".xlsx": - # #152 — binary workbook: metadata only (parity with _scan_vault). - raw = "" - content_preview = "" + # #153 A5 — index sheet names + header rows as text (see _scan_vault). + raw = extract_xlsx_indexable(fpath) + content_preview = raw[:200].strip() else: raw = fpath.read_text(encoding="utf-8", errors="replace") content_preview = raw[:200].strip() diff --git a/backend/search.py b/backend/search.py index ab4072d..b3c23f0 100644 --- a/backend/search.py +++ b/backend/search.py @@ -12,7 +12,12 @@ from sortedcontainers import SortedList from backend import indexer as _indexer from backend import semantic_search as _semantic -from backend.indexer import index + +# NOTE: the shared index is read through ``_indexer.index`` everywhere, never +# via ``from backend.indexer import index``. That import binds the dict object +# once, so a module reload of ``backend.indexer`` (tests, dev reload) rebinds +# the module-level name to a FRESH dict while this module keeps writing to the +# stale one — the inverted index then silently indexes nothing (BUG-089). from backend.services.regex_safety import ( MAX_REGEX_MATCHES, truncate_for_regex, @@ -399,7 +404,7 @@ class InvertedIndex: self.vault_docs = defaultdict(set) self.tag_docs = defaultdict(set) - for vault_name, vault_data in index.items(): + for vault_name, vault_data in _indexer.index.items(): for file_info in vault_data.get("files", []): doc_key = f"{vault_name}::{file_info['path']}" self.doc_count += 1 @@ -688,7 +693,7 @@ _indexer.set_index_change_hook(_on_index_change_hook) def init_inverted_index(): """Force initial inverted index build. Called after build_index completes on startup.""" - if any(vdata.get("files") for vdata in index.values()): + if any(vdata.get("files") for vdata in _indexer.index.values()): _inverted_index.rebuild() logger.info("Inverted index initialized.") @@ -784,7 +789,7 @@ def search( else: candidates = [ (vault_name, file_info) - for vault_name, vault_data in index.items() + for vault_name, vault_data in _indexer.index.items() if vault_filter == "all" or vault_name == vault_filter for file_info in vault_data["files"] ] @@ -1613,7 +1618,7 @@ def get_all_tags(vault_filter: str | None = None) -> dict[str, int]: Dict mapping tag names to their total occurrence count. """ merged: dict[str, int] = {} - for vault_name, vault_data in index.items(): + for vault_name, vault_data in _indexer.index.items(): if vault_filter and vault_filter != "all" and vault_name != vault_filter: continue for tag, count in vault_data.get("tags", {}).items(): diff --git a/backend/services/mutations.py b/backend/services/mutations.py index f2e2aaf..590aec6 100644 --- a/backend/services/mutations.py +++ b/backend/services/mutations.py @@ -19,6 +19,7 @@ import shutil import threading from collections.abc import Callable, Iterator from contextlib import contextmanager +from datetime import date, datetime from pathlib import Path from typing import Any @@ -237,6 +238,14 @@ _XLSX_FLOAT_RE = re.compile(r"^[+-]?(?:\d+\.\d*|\.\d+)$") # the legacy Lotus-style trigger. "+"/"-" are left alone: they are numbers here. _XLSX_FORMULA_RE = re.compile(r"^[=@]") +# #153 A10 — types recognised when a user types into a cell. Excel infers them +# too; storing everything as text would make a spreadsheet unusable (a boolean +# column stays a string, a date column sorts lexicographically). +_XLSX_TRUE_LITERALS = {"true", "vrai", "oui", "yes"} +_XLSX_FALSE_LITERALS = {"false", "faux", "non", "no"} +# Shape check before strptime: keeps the hot path free of format attempts. +_XLSX_DATE_RE = re.compile(r"^\d{1,2}[-/]\d{1,2}[-/]\d{4}(?:[ T]\d{1,2}:\d{2})?$") + # #153 A3 — per-file write lock. Two concurrent saves (two tabs, the AI agent # and the viewer, a watcher restore) would otherwise read-modify-write on the # same archive and the last writer silently wins. Kept deliberately small: the @@ -270,7 +279,21 @@ def _xlsx_write_lock(key: str) -> Iterator[None]: def _coerce_xlsx_value(value: Any) -> Any: - """Turn the string sent by the cell editor back into a scalar.""" + """Turn the string sent by the cell editor back into a scalar (#153 A10). + + The coercion is symmetric with :func:`backend.xlsx_reader._fmt`: a value + typed by the user comes back as a string, and Excel would have inferred a + type when typing the same thing. Recognised here: + + * an empty cell -> ``None`` (clears it) + * ``1234`` / ``-1`` -> ``int`` + * ``1.5`` / ``.5`` -> ``float`` + * ``TRUE``/``FAUX`` (case-insensitive) -> ``bool`` + * ``31/12/2026`` / ``31/12/2026 14:30`` -> ``date``/``datetime`` (FR) + + Anything else stays text. A date-looking string typed with a leading + ``=`` is a formula and never reaches here as a date. + """ if not isinstance(value, str): return value text = value.strip() @@ -280,9 +303,35 @@ def _coerce_xlsx_value(value: Any) -> Any: return int(text) if _XLSX_FLOAT_RE.match(text): return float(text) + lowered = text.lower() + if lowered in _XLSX_TRUE_LITERALS: + return True + if lowered in _XLSX_FALSE_LITERALS: + return False + if not _XLSX_FORMULA_RE.match(text): + parsed = _parse_fr_datetime(text) + if parsed is not None: + return parsed return value +def _parse_fr_datetime(text: str) -> date | datetime | None: + """Parse a FR-localised date/datetime, or return ``None``. + + Accepts ``JJ/MM/AAAA`` and ``JJ/MM/AAAA HH:MM`` (also ``JJ-MM-AAAA``). + ``dayfirst`` is what makes ``01/02/2026`` the 1st of February rather than + the 2nd of January — the French convention. + """ + if not _XLSX_DATE_RE.match(text): + return None + for fmt in ("%d/%m/%Y %H:%M", "%d/%m/%Y", "%d-%m-%Y %H:%M", "%d-%m-%Y"): + try: + return datetime.strptime(text, fmt) + except ValueError: + continue + return None + + def _write_cell(ws: Any, ref: str, value: Any, *, allow_formula: bool) -> None: """Assign one cell, forcing text when it looks like a formula. diff --git a/backend/xlsx_reader.py b/backend/xlsx_reader.py index 4bee9d2..e265b0a 100644 --- a/backend/xlsx_reader.py +++ b/backend/xlsx_reader.py @@ -11,6 +11,7 @@ round-trip would drop (#153 A1) so the UI can warn before saving. from __future__ import annotations import html +import logging import re import zipfile from datetime import date, datetime @@ -20,6 +21,8 @@ from typing import Any from openpyxl import load_workbook from openpyxl.utils import get_column_letter +logger = logging.getLogger("obsigate.xlsx_reader") + # ponytail: hard caps bound the rendered grid (500 rows x 40 cols per sheet). # Raise them, or paginate per sheet, if a real workbook needs more. MAX_ROWS = 500 @@ -48,6 +51,14 @@ _CACHED_FORMULA_RE = re.compile(rb"][^<]*\s*[^<]") # Sheet XML scanned by the cached-formula probe (CPU guard, like MAX_REPLACE_FILE_BYTES). _MAX_PROBE_BYTES = 8_000_000 +# #153 A5 — ceiling on the text handed to the TF-IDF / semantic index. A workbook +# is a data dump, not prose: indexing every cell would flood the inverted index +# and bury the notes. Sheet names + the first rows are enough to make a +# spreadsheet findable by its headers. +MAX_INDEX_CHARS = 5_000 +_INDEX_ROWS_PER_SHEET = 20 +MAX_INDEX_SHEETS = 20 + def _fmt(value: Any) -> str: if value is None: @@ -74,7 +85,29 @@ def _trim(grid: list[list[str]]) -> list[list[str]]: return [row[:width] for row in grid] -def _table(grid: list[list[str]]) -> str: +def _cell_cached(cached: list[list[str]] | None, r: int, c: int) -> str: + """Return the cached result for a 0-based cell, or ``""``. + + The shadow grid is read positionally and may be narrower than the formula + grid (``_trim`` collapses the trailing empty columns of each grid + independently), so every lookup is bounds-checked rather than assumed. + """ + if not cached or r >= len(cached): + return "" + row = cached[r] + return row[c] if c < len(row) else "" + + +def _table(grid: list[list[str]], cached: list[list[str]] | None = None) -> str: + """Render a grid as an HTML table. + + ``cached`` is the same grid read with ``data_only=True`` (#153 A12): where a + formula cell still carries its last computed result, it is shown as a + discreet second line (````) so the user sees the + number Excel last calculated instead of only the formula text. The span + carries ``data-cached-value`` and is titled client-side from + ``xlsx.cached_value_title`` — the backend never emits UI text. + """ if not grid: return "

Feuille vide

" n_cols = max(len(row) for row in grid) @@ -90,7 +123,24 @@ def _table(grid: list[list[str]]) -> str: out.append(f'{r}') for c, val in enumerate(row, start=1): ref = f"{get_column_letter(c)}{r}" - out.append(f'{html.escape(val)}') + # The cached result only makes sense for a formula cell: on a plain + # value cell the two reads are identical and showing both would + # duplicate the text. + shadow = "" + if cached is not None and val.startswith("="): + # `c` is 1-based (A1 notation) and `r` too, while the grid is + # 0-based: translate both. + cval = _cell_cached(cached, r - 1, c - 1) + if cval and cval != val: + # The tooltip is translated client-side from + # `xlsx.cached_value_title`; never hardcode UI text here. + shadow = ( + f'' + f"{html.escape(cval)}" + ) + out.append( + f'{html.escape(val)}{shadow}' + ) out.append("") out.append("") return "".join(out) @@ -142,18 +192,118 @@ def inspect_workbook(file_path: Path) -> list[str]: def render_sheets(file_path: Path) -> list[dict[str, str]]: - """Return ``[{"name": sheet_title, "html": table_html}, ...]``.""" + """Return ``[{"name": sheet_title, "html": table_html}, ...]``. + + Reads the workbook twice: once with ``data_only=False`` for the formulas + (what the user must edit) and, when any formula carries a cached result + (#153 A12), once with ``data_only=True`` to show what Excel last computed. + The second pass is skipped entirely when the archive holds no cached value, + so the common case still costs a single load. + """ wb = load_workbook(str(file_path), read_only=True, data_only=False) try: - sheets = [] - for ws in wb.worksheets: - grid = [ - [_fmt(v) for v in row] - for row in ws.iter_rows( - min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True - ) - ] - sheets.append({"name": ws.title, "html": _table(_trim(grid))}) - return sheets + formulas = [_sheet_grid(ws) for ws in wb.worksheets] + titles = [ws.title for ws in wb.worksheets] finally: wb.close() + + cached: list[list[list[str]]] | None = None + if _has_cached_values(file_path): + cached = _read_cached_grids(file_path, titles) + + sheets = [] + for i, title in enumerate(titles): + grid = _trim(formulas[i]) + # The shadow grid is NOT trimmed independently: _trim drops the + # trailing empty columns of each grid on its own width, which would + # shift every cached value left of its formula. Indexing it + # positionally against the untrimmed grid keeps the two aligned. + shadow = cached[i] if cached is not None and i < len(cached) else None + sheets.append({"name": title, "html": _table(grid, shadow)}) + return sheets + + +def _sheet_grid(ws: Any) -> list[list[str]]: + """Read one worksheet into a grid of formatted strings, bounded by the caps.""" + return [ + [_fmt(v) for v in row] + for row in ws.iter_rows( + min_row=1, max_row=MAX_ROWS, max_col=MAX_COLS, values_only=True + ) + ] + + +def _has_cached_values(file_path: Path) -> bool: + """True when the archive holds at least one ``……``.""" + try: + with zipfile.ZipFile(file_path) as zf: + return _has_cached_formulas(zf) + except (OSError, zipfile.BadZipFile): + return False + + +def _read_cached_grids( + file_path: Path, titles: list[str] +) -> list[list[list[str]]] | None: + """Read every sheet with ``data_only=True`` (what Excel last computed). + + Best effort: returns ``None`` on any failure so the viewer falls back to the + formula-only rendering. A workbook Excel opens but openpyxl cannot re-read + must still display. + """ + try: + wb = load_workbook(str(file_path), read_only=True, data_only=True) + except Exception: + return None + try: + grids = [_sheet_grid(ws) for ws in wb.worksheets] + if [ws.title for ws in wb.worksheets] != titles: + return None + return grids + except Exception: + logger.debug("xlsx cached values unavailable", exc_info=True) + return None + finally: + wb.close() + + +def extract_indexable_text(file_path: Path) -> str: + """Return searchable text for the TF-IDF / semantic index (#153 A5). + + Sheet names plus the first :data:`_INDEX_ROWS_PER_SHEET` rows of each + sheet, capped at :data:`MAX_INDEX_CHARS`. Rows are tab-joined so a search + for a header matches the sheet it belongs to. + + Never raises: a corrupt, encrypted or unsupported workbook yields ``""`` so + the file still gets indexed by name (same contract as :func:`inspect_workbook`). + """ + chunks: list[str] = [] + budget = MAX_INDEX_CHARS + try: + wb = load_workbook(str(file_path), read_only=True, data_only=True) + except Exception: + # Encrypted (BadZipFile) or not a real workbook: name-only indexing. + return "" + try: + for ws in wb.worksheets[:MAX_INDEX_SHEETS]: + if budget <= 0: + break + # The sheet title alone is a strong signal ("Recettes", "Budget"). + block = [ws.title] + for row in ws.iter_rows( + min_row=1, max_row=_INDEX_ROWS_PER_SHEET, max_col=MAX_COLS, values_only=True + ): + cells = [_fmt(v) for v in row] + # Skip blank rows instead of emitting runs of tabs. + if not any(c.strip() for c in cells): + continue + block.append("\t".join(cells).rstrip()) + text = "\n".join(block) + chunks.append(text[:budget]) + budget -= len(text) + except Exception: + # Truncated but still useful: keep whatever was collected. + pass + finally: + wb.close() + return "\n".join(c for c in chunks if c).strip() diff --git a/desktop/Cargo.lock b/desktop/Cargo.lock index b9cf9ae..53b4add 100644 --- a/desktop/Cargo.lock +++ b/desktop/Cargo.lock @@ -2626,7 +2626,7 @@ dependencies = [ [[package]] name = "obsigate-desktop" -version = "2.29.0" +version = "2.30.0" dependencies = [ "chrono", "env_logger", diff --git a/desktop/Cargo.toml b/desktop/Cargo.toml index ad346dc..f40bc51 100644 --- a/desktop/Cargo.toml +++ b/desktop/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "obsigate-desktop" -version = "2.29.0" +version = "2.30.0" description = "ObsiGate Desktop — Porte d'entrée native pour vos vaults Obsidian" authors = ["Bruno Charest"] edition = "2021" diff --git a/desktop/tauri.conf.json b/desktop/tauri.conf.json index 9b21c97..f173fed 100644 --- a/desktop/tauri.conf.json +++ b/desktop/tauri.conf.json @@ -1,7 +1,7 @@ { "$schema": "https://raw.githubusercontent.com/nicedoc/obsigate/main/desktop/tauri.conf.schema.json", "productName": "ObsiGate", - "version": "2.29.0", + "version": "2.30.0", "identifier": "com.obsigate.desktop", "build": { "frontendDist": "../frontend", diff --git a/docs/ISSUES_TODOLIST.md b/docs/ISSUES_TODOLIST.md index 1aef536..655e0d8 100644 --- a/docs/ISSUES_TODOLIST.md +++ b/docs/ISSUES_TODOLIST.md @@ -195,6 +195,7 @@ Avant de corriger quoi que ce soit, un agent IA doit : | *BUG-086* | Édition d'un `.xlsx` : `wb.save()` écrit en place, un plantage laisse un classeur corrompu | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::edit_xlsx_cells` | Simuler un `OSError` pendant `Workbook.save` → le fichier d'origine est tronqué | Écriture atomique : `wb.save(..tmp)` puis `os.replace()` ; `.tmp` supprimé sur échec ; le backup `.bak` reste inchangé. Test : `TestXlsxAtomicWrite::test_failed_save_keeps_the_original` (octets identiques après échec) + `test_no_tmp_left_after_a_successful_save` | #153 A2. Le fichier temporaire a un suffixe `.tmp` → ignoré par le watcher (`_is_relevant` ne retient que les extensions supportées). Vérifié : cf. BUG-085 | | *BUG-087* | Édition d'un `.xlsx` concurrente (deux onglets, agent IA + viewer) : read-modify-write sans verrou, le dernier écrivain gagne silencieusement | 🟢 corrigé | P1 | tableur Excel | IA | `backend/services/mutations.py::_xlsx_write_lock` | Deux `PUT xlsx/save` simultanés sur le même fichier → une écriture est écrasée sans trace | Verrou par chemin (registre + garde, timeout 15 s) autour du cycle load → edit → `os.replace` ; attente dépassée → **409** `conflict`. L'endpoint est devenu `def` (sync) pour que l'attente s'exécute dans le threadpool et ne bloque pas la boucle d'événements. Test : `TestXlsxWriteLock` (2) | #153 A3. Verrou en mémoire, par processus : protège les cas d'un même serveur (le cas desktop/Tauri). Vérifié : cf. BUG-085 | | *BUG-088* | Injection de formule dans un `.xlsx` : une saisie `=cmd\|'/c calc'!A1` est stockée comme formule et s'exécute à l'ouverture dans Excel (DDE) | 🟢 corrigé | P0 | tableur Excel / sécurité | IA | `backend/services/mutations.py::_write_cell`, `backend/routers/files_write.py`, `frontend/js/viewer.js::renderXlsxViewer` | `PUT /api/file/V/xlsx/save` avec `{"sheet": "S", "cells": {"A1": "=1+1"}}` → la cellule sort en `data_type == "f"` | `cell.data_type = "s"` après affectation : le texte est stocké comme chaîne, aucun `` n'est écrit. Opt-in via `allow_formula: true` (endpoint) et le bouton `f(x)` de la visionneuse (session, jamais persisté). Test : `TestXlsxFormulaGuard` (4) + `xlsx-viewer.test.mjs` (toggle) | #153 A4. `+`/`-` ne sont pas neutralisés : ils sont déjà convertis en nombre par `_coerce_xlsx_value`. Le handler global `ServiceError` expose désormais `code` + `details` (le client en a besoin pour le 409), et `api()` (frontend) les propage sur l'Error. Vérifié : cf. BUG-085 | +| *BUG-089* | Un reindex manuel ne reconstruisait pas l'index inversé : la recherche TF-IDF continuait de servir un index périmé | 🟢 corrigé | P1 | ⚙️ backend / recherche | IA | `backend/indexer.py::reload_index`, `backend/indexer.py::reload_single_vault`, `backend/search.py` | Modifier le contenu d'un fichier, puis `GET /api/index/reload` → la recherche renvoie encore l'ancien contenu (ou rien pour un fichier nouveau) | `reload_index()` / `reload_single_vault()` appellent `init_inverted_index()` après le rebuild (le remplacement wholesale d'une entrée de vault n'émet pas les notifications incrémentales). En prime, `backend/search.py` lisait l'index via `from backend.indexer import index` (liaison **par valeur** du dict) : un `importlib.reload(backend.indexer)` recréait le dict côté indexer tandis que la recherche écrivait encore dans l'ancien — l'index inversé n'indexait alors plus rien. Tous les accès passent désormais par `_indexer.index`. Contre-preuve : `TestXlsxSearchable::test_search_finds_a_word_stored_in_a_cell` échoue sans le correctif | #153 A5. Trouvé en écrivant le test de recherche d'A5 : il passait isolément et échouait en suite complète selon l'ordre. Le reload incrémental par fichier (watcher, edition) n'est pas concerné : il passe par le hook `_on_index_change`. Vérifié : suite 1402 passed / 6 skipped, ruff/mypy 0 | ### TODOs techniques (améliorations / nouvelles tâches) @@ -212,6 +213,7 @@ Avant de corriger quoi que ce soit, un agent IA doit : | Date | ID(s) traité(s) | Action | Fichiers modifiés | Résumé | Statut après | |---|---|---|---|---|---| +| 2026-09-28 | BUG-089 (#153 A5, A10, A12) | Correction | `backend/xlsx_reader.py`, `backend/indexer.py`, `backend/search.py`, `backend/services/mutations.py`, `frontend/js/viewer.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `tests/test_xlsx_viewer.py` | **Les tableurs deviennent visibles ettypés** : (A5) `extract_indexable_text()` indexe noms de feuilles + 20 premières lignes (plafond 5 k caractères) dans le TF-IDF et la recherche sémantique — un mot tapé dans une cellule rend le fichier trouvable ; (A10) `_coerce_xlsx_value()` reconnaît désormais les booléens (`TRUE`/`FAUX`/`OUI`/`NON`) et les dates FR `JJ/MM/AAAA` (jour-first : `01/02/2026` = 1er février), symétrique avec l'affichage ; (A12) la valeur calculée en cache s'affiche sous la formule (``, 2ᵉ lecture `data_only=True` uniquement si l'archive contient un ``), info-bulle traduite via `xlsx.cached_value_title` FR/EN. (BUG-089) un reindex manuel reconstruisait mal l'index inversé et `backend/search.py` lisait l'index par valeur. Contre-preuves vérifiées pour A5, A10 et A12. Vérifié : `test_xlsx_viewer.py` 43 passed, suite 1402 passed / 6 skipped, ruff 0, mypy 0, i18n parity, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10 | 🟢 corrigé (en attente vérif utilisateur) | | 2026-09-27 | BUG-085 → BUG-088 (#153 A1-A4) | Correction | `backend/xlsx_reader.py`, `backend/services/mutations.py`, `backend/routers/files_read.py`, `backend/routers/files_write.py`, `backend/schemas.py`, `backend/main.py`, `frontend/js/viewer.js`, `frontend/js/auth.js`, `frontend/style.css`, `frontend/locales/{fr,en}.json`, `frontend/sw.js`, `tests/test_xlsx_viewer.py`, `tests/frontend/xlsx-viewer.test.mjs`, `tests/e2e/xlsx-viewer.spec.js`, `test_vault/sample-xlsx-lossy.xlsx`, `.gitea/workflows/ci.yml` | **Garde-fous d'écriture des classeurs Excel** : (BUG-085) `inspect_workbook()` détecte ce qu'un round-trip openpyxl perd (valeurs calculées, slicers, contrôles, connexions, custom XML, signature) → la lecture expose `xlsx_lossy_features`, la visionneuse affiche une bannière et `PUT xlsx/save` refuse sans `force` (**409** `xlsx_lossy_content`, confirmation explicite puis reprise) ; (BUG-086) écriture atomique `.tmp` + `os.replace` ; (BUG-087) verrou par fichier (409 `conflict`, endpoint sync pour le threadpool) ; (BUG-088) une saisie `=`/`@` est stockée en texte (`data_type = "s"`), sauf opt-in `allow_formula` / bouton `f(x)`. Le handler `ServiceError` expose désormais `code` + `details` et `api()` les propage. Périmètre de perte revalidé empiriquement sur openpyxl 3.1.5 (graphiques, images et TCD sont préservés). Vérifié : `test_xlsx_viewer.py` 31 passed, suite 1390 passed / 6 skipped, ruff/mypy 0, validate-imports 40 modules, xlsx-viewer.test.mjs 10/10, E2E 3/3 | 🟢 corrigé (en attente vérif utilisateur) | | *(exemple)* 2026-06-15 | BUG-001 | Correction | `frontend/app.js` | Réécriture de `renderFile()` pour préserver le DOM dashboard | 🟢 corrigé (en attente vérif) | | 2026-09-09 | BUG-001, BUG-002 | Correction | `backend/main.py`, `frontend/excalidraw-editor.html`, `tests/test_pdf_stream.py` | BUG-001: Content-Disposition RFC 5987 (nom PDF accentué ne casse plus l'en-tête → plus de 500). BUG-002: suppression alias esm.sh (408 jotai) + React 19 cohérent + prop `excalidrawAPI` → Loading masqué, save OK. Vérifié: 534 tests backend verts + E2E navigateur. | 🟢 corrigé (en attente vérif utilisateur) | diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 66e5bec..e009a43 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -1,6 +1,6 @@ # ObsiGate — Roadmap -> **Version :** 2.29.0 | **Dernière mise à jour :** 2026-09-27 +> **Version :** 2.30.0 | **Dernière mise à jour :** 2026-09-27 > **Ce fichier ne contient que le travail à venir** (🔵 En cours + ⚪ Backlog) et un index compact > vers les fonctionnalités livrées. > - **Méthode de livraison à appliquer pour toute tâche : [DELIVERY_WORKFLOW.md](./DELIVERY_WORKFLOW.md)** @@ -47,7 +47,7 @@ ### 153. Visionneuse & édition XLSX — complétude (fidélité, recherche, IA, UX, formats) - **Effort :** 8-13 jours (P0 ✅ 2-3 j · P1 : 4-6 j · P2 : 2-4 j) | **Impact :** 🟡 -- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), reste P1 puis P2 +- **Statut :** 🔵 en cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17 - **Analyse, risques et critères d'acceptation :** [features/xlsx-viewer.md](./features/xlsx-viewer.md) - **Description :** #152 (visionneuse XLSX, 2.27.0) lit et édite correctement la **grille de valeurs** d'un `.xlsx`, mais l'ensemble supporté est étroit : valeurs seulement (ni structure, @@ -70,15 +70,15 @@ - [x] **A2** Écriture atomique (`wb.save(.tmp)` + `os.replace()`, backup inchangé) — BUG-086 - [x] **A3** Verrou par fichier autour du read-modify-write (timeout 15 s + **409** `conflict`) — BUG-087 - [x] **A4** Neutralisation de l'injection de formule (`=`/`@` stockés en texte, opt-in `allow_formula` + bouton `f(x)`) — BUG-088 - - **P1 — recherche, IA, UX (🟡, 4-6 j) — ⚪ à faire** - - [ ] **A5** Indexation du contenu des feuilles (TF-IDF + sémantique, plafond ~5 k caractères) + - **P1 — recherche, IA, UX (🟡, 4-6 j) — 🔵 en cours** + - [x] **A5** Indexation du contenu des feuilles (noms de feuilles + 20 premières lignes, plafond 5 k caractères) — les mots tapés dans une cellule rendent le fichier trouvable ; au passage **BUG-089** (reindex manuel ne reconstruisait pas l'index inversé) - [ ] **A6** Outils IA `update_xlsx_cells` / `append_xlsx_rows` / `xlsx_to_markdown` / `list_xlsx_sheets` - [ ] **A7** Navigation clavier + barre de formule + nom de cellule (Tab/Entrée/flèches, `Maj+Entrée`, copie de plage) - [ ] **A8** `thead` sticky + bandeau « feuille tronquée » (lève la troncature silencieuse) - [ ] **A9** Chargement paresseux par feuille (`GET …/xlsx/sheet?offset&limit`, défilement virtuel) - - [ ] **A10** Types & formats de saisie (nombre/texte, booléens, dates localisées FR) + - [x] **A10** Types & formats de saisie (nombre/texte, booléens `TRUE`/`FAUX`, dates FR `JJ/MM/AAAA` jour-first) - [ ] **A11** Tests frontend (`tests/frontend/xlsx-viewer.test.mjs`) + E2E (`tests/e2e/xlsx-viewer.spec.js`) au CI - - [ ] **A12** Affichage de la valeur calculée en cache (lecture `data_only=True`, mention FR/EN) + - [x] **A12** Valeur calculée affichée sous la formule (2ᵉ lecture `data_only=True` seulement si l'archive contient un ``, info-bulle FR/EN) - **P2 — étendu (🟢, 2-4 j) — ⚪ à faire** - [ ] **A13** Tri / filtre / recherche dans la feuille + export CSV de la sélection - [ ] **A14** CRUD de feuilles, lignes et colonnes (renommer, insérer, supprimer, dupliquer) diff --git a/docs/features/xlsx-viewer.md b/docs/features/xlsx-viewer.md index 8acef42..7e5f1c9 100644 --- a/docs/features/xlsx-viewer.md +++ b/docs/features/xlsx-viewer.md @@ -3,7 +3,7 @@ > **Item de roadmap :** [#153 — Visionneuse & édition XLSX — complétude](../ROADMAP.md) > **Origine :** #152 (visionneuse XLSX, livrée en 2.27.0 — voir > [archive/COMPLETED_v1-v2.md](../archive/COMPLETED_v1-v2.md)) -> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), P1/P2 restants +> **Statut :** 🔵 En cours — **P0 livré le 2026-09-27** (BUG-085 → BUG-088), **A5/A10/A12 livrés le 2026-09-28** (avec BUG-089), reste A6-A9 puis A13-A17 > **Effort estimé :** 8-13 jours au total (P0 ✅ 2-3 j · P1 4-6 j · P2 2-4 j) > **Règle de maintenance :** la Roadmap porte les cases à cocher (suivi), cette fiche porte > l'analyse, les risques et les critères d'acceptation. **Ne pas dupliquer le détail.** @@ -134,13 +134,14 @@ couverture) · effort en jours-homme de développement + tests. nombres. Au passage : le handler `ServiceError` expose `code` + `details` et `api()` les propage sur l'Error. *Vérifié :* `TestXlsxFormulaGuard` (4) + test du toggle côté UI. -### P1 — Recherche, IA, UX (4-6 j) +### P1 — Recherche, IA, UX (4-6 j) — 🟢 A5, A10, A12 livrés le 2026-09-28 -- [ ] **A5 — Indexation du contenu des feuilles.** Extraire un texte (noms de feuilles + - en-têtes + N premières lignes, plafond ~5 k caractères) pour le TF-IDF et la recherche - sémantique, tout en gardant la lecture binaire pour l'affichage ; `content_preview` - renseigné ; exclusion si le classeur est chiffré/corrompu. *Critère :* une cellule contenant - un mot-clé rend le fichier trouvable ; `test_xlsx_indexing` étendu. +- [x] **A5 — Indexation du contenu des feuilles.** `extract_indexable_text()` (noms de feuilles + + 20 premières lignes, `MAX_INDEX_CHARS = 5 000`, 20 feuilles max) alimente le TF-IDF et la + recherche sémantique ; la lecture binaire reste inchangée pour l'affichage. Un classeur + chiffré/corrompu s'indexe par son seul nom (jamais d'exception). Au passage : **BUG-089**, + un reindex manuel ne reconstruisait pas l'index inversé. *Vérifié :* `TestXlsxSearchable` (4) + + `TestXlsxIndexing`, **contre-preuve** (neutraliser l'extraction → 3 tests échouent). - [ ] **A6 — Outils IA sur classeur.** `update_xlsx_cells` (enveloppe du service existant), `append_xlsx_rows`, `xlsx_to_markdown` (contexte LLM, plafonné), `list_xlsx_sheets` — risque WRITE + confirmation pour les mutations, libellés i18n dans `backend/tools/labels.py`, @@ -155,15 +156,22 @@ couverture) · effort en jours-homme de développement + tests. `GET /api/file/{vault}/xlsx/sheet?sheet=N&offset=&limit=` (`response_model` + `backend/openapi_docs.py`), rendu à la demande avec défilement virtuel, bouton « charger tout ». -- [ ] **A10 — Types et formats de saisie.** Coercion symétrique à l'écriture/à l'affichage - (nombre vs texte, booléens `TRUE`/`FALSE`, dates localisées FR — le TODO existe déjà dans - `_coerce_xlsx_value`) ; affichage du type d'origine dans l'info-bulle de cellule. +- [x] **A10 — Types et formats de saisie.** `_coerce_xlsx_value()` reconnait les booléens + (`true`/`vrai`/`oui`/`yes` et leurs négatifs) et les dates FR `JJ/MM/AAAA` (+ `HH:MM`), + jour-first comme Excel en locale française : `01/02/2026` = 1ᵉʳ février. Une saisie + ressemblant à une formule n'est jamais convertie (BUG-088 préservé) ; un code postal + numérique ou une version restent ce qu'ils sont. *Vérifié :* `TestXlsxValueCoercion` (5), + **contre-preuve** (neutraliser la coercion → 2 tests échouent). - [ ] **A11 — Tests frontend + E2E.** `tests/frontend/xlsx-viewer.test.mjs` (dirty, Échap, collage, 1 PUT par feuille, bouton désactivé) et `tests/e2e/xlsx-viewer.spec.js` (ouverture, onglets, édition, sauvegarde, rechargement) ; intégration au CI. -- [ ] **A12 — Valeurs calculées.** Afficher la valeur en cache (2ᵉ ligne discrète) quand elle - existe, via une lecture `data_only=True` de la même page d'onglets ; mention FR/EN - « valeur recalculée par Excel ». +- [x] **A12 — Valeurs calculées.** La valeur en cache s'affiche sous la formule dans un + ``. La 2ᵉ lecture `data_only=True` n'a lieu que si l'archive + contient réellement un `……` (sonde déjà présente pour A1) : le cas courant + reste à un seul chargement, et toute erreur retombe sur l'affichage formules seul. + L'info-bulle est traduite côté client (`xlsx.cached_value_title` FR/EN) — aucun texte + d'interface n'est émis par le backend. *Vérifié :* `TestXlsxCachedValues` (3), + **contre-preuve** (neutraliser la 2ᵉ lecture → 2 tests échouent). ### P2 — Étendu (2-4 j) @@ -199,3 +207,4 @@ couverture) · effort en jours-homme de développement + tests. | 2026-09-27 | Audit complet → création de #153 : limites, risques R1-R5, backlog A1-A17 | | 2026-09-27 | Périmètre de perte **remesuré** sur openpyxl 3.1.5 : graphiques / images / TCD sont préservés, seules les valeurs en cache et quelques parties exotiques sont perdues | | 2026-09-27 | **P0 livré** (BUG-085 → BUG-088) : `xlsx_lossy_features` + 409 `xlsx_lossy_content`, écriture atomique, verrou par fichier, formules stockées en texte par défaut | +| 2026-09-28 | **A5 + A10 + A12 livrés** : le contenu des cellules est indexé (recherche), la saisie est typée (booléens, dates FR), la valeur calculée s'affiche sous la formule. **BUG-089** corrigé au passage (reindex manuel ≠ reconstruction de l'index inversé ; `backend/search.py` lisait l'index par valeur) | diff --git a/frontend/js/viewer.js b/frontend/js/viewer.js index ea36747..3cae602 100644 --- a/frontend/js/viewer.js +++ b/frontend/js/viewer.js @@ -1056,6 +1056,12 @@ export function renderXlsxViewer(area, data) { const dirtyCount = () => area.querySelectorAll("td.xlsx-dirty").length; const refreshSaveState = () => { saveBtn.disabled = dirtyCount() === 0; }; + // #153 A12 — the backend marks the last value Excel computed; the wording is + // translated here so the tooltip follows the UI language. + area.querySelectorAll(".xlsx-cached[data-cached-value]").forEach((el) => { + el.title = t("xlsx.cached_value_title"); + }); + // Editable cells: Enter blurs, Escape reverts, paste stays single-line. area.querySelectorAll(".xlsx-table td").forEach((td) => { td.contentEditable = "true"; diff --git a/frontend/locales/en.json b/frontend/locales/en.json index a803c72..13ab19a 100644 --- a/frontend/locales/en.json +++ b/frontend/locales/en.json @@ -1828,6 +1828,7 @@ "xlsx.lossy_confirm": "Save anyway? The following will be lost: {features}", "xlsx.lossy_cancelled": "Save cancelled", "xlsx.formula_toggle_title": "Treat “=” and “@” as formulas (off by default)", + "xlsx.cached_value_title": "Last value calculated by Excel", "xlsx.feature_cached_values": "cached values", "xlsx.feature_slicers": "slicers and timelines", "xlsx.feature_form_controls": "form controls", diff --git a/frontend/locales/fr.json b/frontend/locales/fr.json index 383547b..140bc8b 100644 --- a/frontend/locales/fr.json +++ b/frontend/locales/fr.json @@ -1828,6 +1828,7 @@ "xlsx.lossy_confirm": "Enregistrer quand même ? Les éléments suivants seront perdus : {features}", "xlsx.lossy_cancelled": "Sauvegarde annulée", "xlsx.formula_toggle_title": "Interpréter « = » et « @ » comme des formules (désactivé par défaut)", + "xlsx.cached_value_title": "Dernière valeur calculée par Excel", "xlsx.feature_cached_values": "valeurs calculées", "xlsx.feature_slicers": "segments et chronologies", "xlsx.feature_form_controls": "contrôles de formulaire", diff --git a/frontend/style.css b/frontend/style.css index 5f87302..c153b08 100644 --- a/frontend/style.css +++ b/frontend/style.css @@ -10987,6 +10987,20 @@ body.desktop-mode .editor-container { background: rgba(255, 196, 0, 0.18); } +/* #153 A12 — last result Excel computed, shown under a formula cell. + Discreet by design: the formula is what the user edits, the cached value is + context (stale until Excel recalculates). */ +.xlsx-cached { + display: block; + margin-top: 2px; + padding-left: 6px; + border-left: 2px solid var(--border, #d0d7de); + color: var(--text-muted); + font-size: 0.85em; + font-variant-numeric: tabular-nums; + white-space: nowrap; +} + /* #153 A1/A4 — lossy-save warning + formula toggle */ .xlsx-warning { display: flex; diff --git a/package.json b/package.json index 5c7257e..ac4c8e2 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "obsigate", - "version": "2.29.0", + "version": "2.30.0", "description": "**Porte d'entrée web ultra-léger pour vos vaults Obsidian** — Accédez, naviguez et recherchez dans toutes vos notes Obsidian depuis n'importe quel appareil via une interface web moderne et responsive.", "main": "patch.js", "directories": { diff --git a/tests/e2e/xlsx-viewer.spec.js b/tests/e2e/xlsx-viewer.spec.js index a7322bd..69fa24e 100644 --- a/tests/e2e/xlsx-viewer.spec.js +++ b/tests/e2e/xlsx-viewer.spec.js @@ -11,7 +11,8 @@ * - the warning banner lists both features ; * - saving a cell on that workbook asks for confirmation (native dialog) and * then succeeds (the client retries with `force: true`) ; - * - the f(x) toggle is off by default, so "=B1*3" is stored as text. + * - the f(x) toggle is off by default, so "=B1*3" is stored as text ; + * - the value Excel last computed is shown under the formula (#153 A12). * * The fixture is restored byte-for-byte in `afterAll` so a local run never * dirties the working copy. @@ -53,7 +54,7 @@ async function openFixture(page) { await expect(page.locator('#content-area .xlsx-table')).toBeVisible({ timeout: 15000 }); } -test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => { +test.describe('Excel viewer — garde-fous d\'écriture et valeurs calculées (#153)', () => { test.beforeAll(() => { if (existsSync(FIXTURE_PATH)) originalBytes = readFileSync(FIXTURE_PATH); }); @@ -62,6 +63,28 @@ test.describe('Excel viewer — garde-fous d\'écriture (#153 P0)', () => { if (originalBytes) writeFileSync(FIXTURE_PATH, originalBytes); }); + // Read-only assertions come FIRST, before the mutating tests: saving through + // the viewer rewrites the workbook and an openpyxl round-trip drops the cached + // formula results (BUG-085), so the shadow line only exists on a pristine + // fixture. + test('affiche la valeur calculée en cache sous la formule (#153 A12)', async ({ page }) => { + await login(page); + await openFixture(page); + + // B2 is "=B1*2" and the package keeps its last result (200). + const formulaCell = page.locator('#content-area td[data-cell="B2"]'); + await expect(formulaCell).toContainText('=B1*2'); + + const cached = formulaCell.locator('.xlsx-cached'); + await expect(cached).toHaveCount(1); + await expect(cached).toHaveText('200'); + // The tooltip is translated client-side, never hardcoded by the backend. + await expect(cached).toHaveAttribute('title', /Excel/); + + // A plain value cell must not be duplicated with a shadow line. + await expect(page.locator('#content-area td[data-cell="B1"] .xlsx-cached')).toHaveCount(0); + }); + test('affiche la bannière listant les éléments non préservés', async ({ page }) => { await login(page); await openFixture(page); diff --git a/tests/test_xlsx_viewer.py b/tests/test_xlsx_viewer.py index 1fd78ed..b833523 100644 --- a/tests/test_xlsx_viewer.py +++ b/tests/test_xlsx_viewer.py @@ -106,16 +106,208 @@ class TestXlsxIndexing: assert ".xlsx" in SUPPORTED_EXTENSIONS - def test_xlsx_indexed_metadata_only(self, test_vault_dir, xlsx_file): + def test_xlsx_indexes_sheet_names_and_headers(self, test_vault_dir, xlsx_file): + """#153 A5 — a workbook is searchable by its cell values.""" from backend.indexer import _index_single_file_sync info = _index_single_file_sync(VAULT, test_vault_dir, xlsx_file) assert info is not None assert info["extension"] == ".xlsx" - assert info["content"] == "" # binary: never read into TF-IDF + # Sheet names + header rows reach TF-IDF (was metadata-only before A5). + assert "Budget" in info["content"] + assert "Poste" in info["content"] + assert info["content_preview"] assert info["title"] # filename-derived title +class TestXlsxCachedValues: + """#153 A12 — show what Excel last computed next to each formula.""" + + def test_cached_result_is_shown_beside_the_formula(self, client, lossy_xlsx): + from backend.xlsx_reader import render_sheets + + html = render_sheets(Path(lossy_xlsx))[0]["html"] + # A2 is "=A1*3" with 9 in the fixture. + assert 'data-cell="A2"' in html + assert "=A1*3" in html + assert "xlsx-cached" in html + assert ">9<" in html # the cached result Excel computed + + def test_no_shadow_when_no_formula_carries_a_result(self, client, xlsx_file): + from backend.xlsx_reader import render_sheets + + html = render_sheets(Path(xlsx_file))[0]["html"] + assert "xlsx-cached" not in html # budget.xlsx has no at all + + def test_plain_cells_are_never_duplicated(self, client, lossy_xlsx): + from backend.xlsx_reader import render_sheets + + html = render_sheets(Path(lossy_xlsx))[0]["html"] + # A1 is the literal 3: the two reads agree, so only one value shows. + assert 'data-cell="A1"' in html + assert html.count("xlsx-cached") == 1 # only the formula cell + + +class TestXlsxValueCoercion: + """#153 A10 — a typed value comes back with the type Excel would infer.""" + + def _write(self, client, xlsx_file, ref, value): + return client.put( + f"/api/file/{VAULT}/xlsx/save", + params={"path": "budget.xlsx"}, + json={"sheet": "Budget", "cells": {ref: value}, "force": True}, + ) + + def test_number_and_bool_are_stored_as_typed(self, client, xlsx_file): + resp = self._write( + client, xlsx_file, "D1", "42" + ) + assert resp.status_code == 200 + resp = self._write(client, xlsx_file, "D2", "VRAI") + assert resp.status_code == 200 + resp = self._write(client, xlsx_file, "D3", "12/03/2026") + assert resp.status_code == 200 + + from openpyxl import load_workbook + + wb = load_workbook(xlsx_file) + ws = wb["Budget"] + assert ws["D1"].value == 42 and isinstance(ws["D1"].value, int) + assert ws["D2"].value is True + assert ws["D3"].value.year == 2026 and ws["D3"].value.month == 3 + assert ws["D3"].value.day == 12 # FR day-first, not 3 December + wb.close() + + def test_day_first_date_is_not_read_as_us(self, client, xlsx_file): + """'01/02/2026' is 1 February in French, not 2 January.""" + assert self._write(client, xlsx_file, "E1", "01/02/2026").status_code == 200 + from openpyxl import load_workbook + + wb = load_workbook(xlsx_file) + d = wb["Budget"]["E1"].value + wb.close() + assert (d.month, d.day) == (2, 1) + + def test_ambiguous_text_is_left_alone(self, client, xlsx_file): + """A version string or a partial date stays text, never a date.""" + assert self._write(client, xlsx_file, "F1", "3.14.2").status_code == 200 + assert self._write(client, xlsx_file, "F2", "Ref 12/34").status_code == 200 + assert self._write(client, xlsx_file, "F3", "12/2026").status_code == 200 + from openpyxl import load_workbook + + wb = load_workbook(xlsx_file) + ws = wb["Budget"] + assert ws["F1"].value == "3.14.2" + assert ws["F2"].value == "Ref 12/34" + assert ws["F3"].value == "12/2026" + wb.close() + + def test_plain_integer_text_becomes_a_number(self, client, xlsx_file): + """Typing a bare number yields a number, as it did before #153 A10.""" + assert self._write(client, xlsx_file, "H1", "75001").status_code == 200 + from openpyxl import load_workbook + + wb = load_workbook(xlsx_file) + assert wb["Budget"]["H1"].value == 75001 + wb.close() + + def test_formula_looking_date_stays_text(self, client, xlsx_file): + """A formula is never mistaken for a date (BUG-088 must not regress).""" + assert self._write(client, xlsx_file, "G1", "=12/03/2026").status_code == 200 + from openpyxl import load_workbook + + wb = load_workbook(xlsx_file) + c = wb["Budget"]["G1"] + wb.close() + assert c.data_type == "s" + assert c.value == "=12/03/2026" + + +class TestXlsxSearchable: + """#153 A5 — a keyword living in a CELL must make the file findable.""" + + def test_search_finds_a_word_stored_in_a_cell(self, test_vault_dir): + """A word typed in a CELL must make the workbook findable. + + Deliberately hermetic: it drives the indexer and the inverted index + directly instead of going through the HTTP reload, because both are + process-wide singletons that other test modules mutate (some reload the + module outright), which would make this test order-dependent. + """ + from openpyxl import Workbook + + import backend.indexer as indexer + import backend.search as search_mod + + path = Path(test_vault_dir) / "fournisseurs.xlsx" + wb = Workbook() + ws = wb.active + ws.title = "Contacts" + ws["A1"] = "Fournisseur" + ws["A2"] = "Menuiserie Beaulieu" + wb.save(path) + + # Index exactly this vault, from scratch. + indexer.index[VAULT] = { + "files": [indexer._index_single_file_sync(VAULT, test_vault_dir, str(path))], + "tags": {}, + "paths": [], + "config": {}, + } + entries = [f for f in indexer.index[VAULT]["files"] if f] + assert entries, "le tableur n'a pas ete indexe" + assert "Menuiserie" in entries[0]["content"], ( + f"contenu indexe : {entries[0]['content'][:80]!r}" + ) + + search_mod.init_inverted_index() + hits = [r["path"] for r in search_mod.search("Menuiserie", "all")] + assert "fournisseurs.xlsx" in hits + + def test_indexable_text_is_capped(self, tmp_path): + """A data dump must not flood the index.""" + from openpyxl import Workbook + + from backend.xlsx_reader import MAX_INDEX_CHARS, extract_indexable_text + + path = tmp_path / "huge.xlsx" + wb = Workbook() + ws = wb.active + ws.title = "Big" + for r in range(1, 400): + ws.cell(row=r, column=1, value=f"ligne {r} " + "x" * 60) + wb.save(path) + + text = extract_indexable_text(path) + assert len(text) <= MAX_INDEX_CHARS + assert "Big" in text # the sheet name survives + + def test_corrupt_workbook_indexes_as_empty_not_a_crash(self, tmp_path): + from backend.xlsx_reader import extract_indexable_text + + path = tmp_path / "broken.xlsx" + path.write_bytes(b"not a zip at all") + assert extract_indexable_text(path) == "" + + def test_blank_rows_are_skipped(self, tmp_path): + from openpyxl import Workbook + + from backend.xlsx_reader import extract_indexable_text + + path = tmp_path / "sparse.xlsx" + wb = Workbook() + ws = wb.active + ws.title = "S" + ws["A1"] = "Alpha" + ws["A50"] = "Omega" # beyond the indexed prefix -> ignored on purpose + wb.save(path) + + text = extract_indexable_text(path) + assert "Alpha" in text + assert "Omega" not in text + assert "\t\t" not in text # no run of empty columns + + # ── Edit ──────────────────────────────────────────────────────────────────