fix(editor+agent): multilignes conservees (cause reelle) + validation des modeles Nvidia
FlowDeck CI / test (push) Failing after 14s
FlowDeck CI / docker (push) Skipped

Editeur:
- sync() lisait el.textContent qui supprime les <br> : ajout de gt()
  (innerText) utilise par sync(), tableaux (toggle/columns) et _resultText.
- Scission sur Entree + limites haut/bas de bloc via splitCaret() (marqueur
  temporaire au curseur), plus de perte de fin de bloc multi-lignes.

Agent:
- fetch_provider_models() valide la liste nvidia : filtre des modeles
  non-chat (nom) puis probe reelle /chat/completions (6 paralleles, 6s),
  seuls les modeles 2xx restent. 429 conserve (route, donc utilisable).
- Repli sur la liste filtree si toutes les validations echouent.
- tests/test_llm_config.py : 5 cas (catalogue mock, 404 chat, autres
  providers non valides, chute de securite, offline). 276 tests passent.
This commit is contained in:
2026-09-07 13:53:43 -04:00
parent 6356cd5703
commit b15289b4bc
6 changed files with 269 additions and 10 deletions
+33
View File
@@ -1,5 +1,38 @@
# Changelog — FlowDeck
## v5.1.8 (2026-09-07) — Multi-lignes réellement conservées + modèles IA validés (Nvidia)
> Deux correctifs : le bug multi-lignes (v5.1.7 corrigeait le symptôme au
> re-rendu, mais la vraie cause était la lecture en `textContent` qui élimine
> les `<br>`) et, côté Agent, seuls les modèles réellement utilisables du
> fournisseur Nvidia sont désormais proposés dans la configuration IA.
### Éditeur — multilignes conservées (vraie cause)
- `sync()` lisait `el.textContent`, qui **supprime** les `<br>` : les sauts de
ligne étaient perdus dès la sauvegarde du bloc. Ajout de `gt()` (basé sur
`innerText`, qui rend bien les `<br>` en `\n`) utilisé par `sync()`, les
tableaux (toggle/columns) et le calcul de position du curseur.
- La scission sur Entrée (et la limite bas/haut de bloc) utilisait désormais
`splitCaret()` : un marqueur temporaire est inséré au curseur puis
`gt()`/`innerText` découpe le contenu en « avant » / « après », exactement
comme un vrai clavier. Fini le décalage `before` qui perdait la fin d'un
bloc multi-lignes.
- `_resultText` (composer IA) lit par `gt()` pour préserver les `\n`.
### Agent — validation des modèles (fournisseur bruité)
- `GET {base}/models` de Nvidia liste tout le catalogue, mais la plupart des
modèles (embeddings, rerank, génération d'images…) répondent **404** sur
`/v1/chat/completions` — l'erreur observée dans la config IA.
- `fetch_provider_models()` filtre désormais pour `nvidia` : suppression des
modèles manifestement non-chat (nom), puis « probe » réelle de chaque
candidat (`POST /chat/completions`, payload minimal, pipeline limité à 6 en
parallèle, 6 s par appel) — seuls les modèles qui répondent sont listés.
Un 429 (limite de débit) est conservé (le modèle est routé, donc utilisable).
- Filet de sécurité : si toutes les validations échouent, on garde la liste
filtrée par nom plutôt que de vider le sélecteur.
- Tests : `tests/test_llm_config.py` (5 cas) — catalogue mock, 404 chat,
retour sans validation pour les autres providers, chute de sécurité, offline.
## v5.1.7 (2026-09-07) — Retours à la ligne conservés quand on change de bloc (Entrée)
> Un bloc multi-lignes (Shift+Entrée) perdait ses retours à la ligne dès qu'on
+1 -1
View File
@@ -1 +1 @@
5.1.7
5.1.8
+1 -1
View File
@@ -60,7 +60,7 @@ async def lifespan(_app: FastAPI):
app = FastAPI(
title="FlowDeck",
version="5.1.7",
version="5.1.8",
docs_url="/docs" if settings.log_level == "DEBUG" else None,
redoc_url=None,
lifespan=lifespan,
+77 -1
View File
@@ -11,11 +11,32 @@ Rows are created lazily, so .env stays the default until saved from the UI.
from __future__ import annotations
import json
import time
from typing import Optional
from app.config import settings
from app.services.llm_client import PROVIDERS, PROVIDER_MODELS
# Providers whose /v1/models lists far more entries than /v1/chat/completions
# actually serves. The fetched list is validated (name filter + live probe)
# before being exposed as "usable models". NVIDIA exposes all of its catalog
# (embeddings, rerank, image/video/audio gen…) many of which answer 404 on
# chat completions — the exact failure the user hit.
_CHAT_VALIDATED_PROVIDERS = frozenset({"nvidia"})
# Markers that identify clearly non-chat models (embeddings, rerank, media gen…).
_NON_CHAT_MARKERS = (
"embed", "bge-", "rerank", "retriev", "tts", "asr", "stt", "whisper", "speech",
"transcrib", "translate", "image", "video", "audio", "music", "sound", "dall",
"stable", "diffus", "flux", "sora", "veo", "midjourney", "clip", "segmentation",
"ocr", "inpainting", "depth", "motion", "sento-",
)
def _is_likely_chat(model_id: str) -> bool:
ml = model_id.lower()
return not any(m in ml for m in _NON_CHAT_MARKERS)
__all__ = [
"get_llm_config", "set_llm_config", "provider_info",
"get_user_llm_key", "list_user_llm_keys", "upsert_user_llm_key",
@@ -315,4 +336,59 @@ async def fetch_provider_models(provider: str, *, api_key: str = "",
if i not in seen:
seen.add(i)
out.append(i)
return out[:300]
# Providers like NVIDIA list their whole catalog, most of which is NOT served
# by /v1/chat/completions (404 sur « model not found »). Filter by name first,
# then probe the survivors with a minimal chat call so only usable models stay.
if provider in _CHAT_VALIDATED_PROVIDERS and out:
candidates = [m for m in out if _is_likely_chat(m)] or out
validated = await _validate_chat_models(base_url, api_key, candidates)
# Ne vidons jamais la liste : en cas d'échec de validation (débit limité,
# indisponibilité passagère) on garde la liste filtrée par nom.
out = validated if validated else candidates
return out[:300]
async def _validate_chat_models(base_url: str, api_key: str, candidates: list[str],
*, timeout: float = 6.0, concurrency: int = 6,
deadline: float = 55.0) -> list[str]:
"""Probe `POST {base_url}/chat/completions` for each candidate with a minimal
payload and keep the models that accept it. A 429 (rate limit) is treated as
"probably fine" since it proves the model is routed, not that it's invalid."""
import asyncio
import httpx
valid: list[str] = []
sem = asyncio.Semaphore(concurrency)
start = time.monotonic()
headers = {"Content-Type": "application/json"}
if api_key:
headers["Authorization"] = f"Bearer {api_key}"
payload: dict = {
"model": None,
"messages": [{"role": "user", "content": "ping"}],
"max_tokens": 4,
}
async def probe(model: str) -> str | None:
if time.monotonic() - start > deadline:
return None
payload["model"] = model
try:
async with sem:
async with httpx.AsyncClient(timeout=timeout) as client:
resp = await client.post(
f"{base_url}/chat/completions", headers=headers, json=payload
)
if resp.status_code < 300 or resp.status_code == 429:
return model
except Exception: # noqa: BLE001 — réseau/timeout ⇒ rejeté
return None
return None
results = await asyncio.gather(*(probe(m) for m in candidates))
for m in results:
if m:
valid.append(m)
return valid
+18 -7
View File
@@ -15,6 +15,16 @@
let _bid=0;function genId(){return 'b'+(++_bid)+'_'+Date.now().toString(36);}
function cp(e){const s=window.getSelection();if(!s.rangeCount)return 0;const r=s.getRangeAt(0).cloneRange();r.selectNodeContents(e);r.setEnd(s.getRangeAt(0).endContainer,s.getRangeAt(0).endOffset);return r.toString().length;}
function gt(e){return (e&&e.innerText!==undefined)?e.innerText:(e?e.textContent:'');}
function splitCaret(e){
const s=window.getSelection(); if(!s.rangeCount||!s.rangeCount)return null; const sel=s.getRangeAt(0);
const MK='\uF000FD'; const mark=document.createElement('span'); mark.textContent=MK;
sel.insertNode(mark);
const full=gt(e)||'';
const parts=full.split(MK);
mark.remove(); e.normalize();
return {before:parts[0]||'',after:(parts.length>1?parts[1]:'')||'',content:full.split(MK).join('')};
}
function ce(e){const r=document.createRange();r.selectNodeContents(e);r.collapse(false);window.getSelection().removeAllRanges();window.getSelection().addRange(r);}
function cs(e){const r=document.createRange();r.selectNodeContents(e);r.collapse(true);window.getSelection().removeAllRanges();window.getSelection().addRange(r);}
function esc(s){return s.replace(/&/g,'&amp;').replace(/</g,'&lt;').replace(/>/g,'&gt;');}
@@ -339,7 +349,7 @@
_err(msg){ this._html='<p style="color:#e5484d">⚠ '+esc(msg)+'</p>'; this._stage='result'; this.paint(); },
insert(){ const E=window.E; if(!E||typeof E.applyAIBlocks!=='function')return; E.sync();
const result=this._resultText(); if(!result)return; E.applyAIBlocks(result); E.dirty=true; E.autoSave(); this.close(); },
_resultText(){ const el=this.ROOT&&this.ROOT.querySelector('.aici-result'); return el?el.textContent:''; },
_resultText(){ const el=this.ROOT&&this.ROOT.querySelector('.aici-result'); if(!el)return ''; const r=document.createRange();r.selectNodeContents(el);return r.toString(); },
toast(m){ if(window.showToast)window.showToast(m,'info'); }
};
window.AIC=AIC;
@@ -741,9 +751,9 @@
sync(){const ct=document.getElementById('_blocksCt');if(!ct)return;for(const b of this.blocks){if(b.type==='table'){const wrap=ct.querySelector(`[data-bid="${b.id}"].ftable-editor`);if(wrap){const rows=[];wrap.querySelectorAll('tr.ftable-row').forEach(function(tr){const cells=tr.querySelectorAll('th.ftable-cell,td.ftable-cell');if(!cells.length)return;const rowarr=[];cells.forEach(function(cell){rowarr.push(cell.textContent||'');});rows.push(rowarr);});if(rows.length)b.rows=rows;}continue;}
if(b.type==='divider'||b.type==='image'||b.type==='embed'||b.type==='table_of_contents')continue;
if(b.type==='math'){const te=ct.querySelector(`[data-math-bid="${b.id}"]`);if(te)b.content=te.value||'';continue;}
const el=ct.querySelector(`[data-bid="${b.id}"]`);if(el&&!el.classList.contains('toggle-title'))b.content=el.textContent||'';else if(el)b.content=el.textContent||'';
if(b.type==='toggle'&&b.children)b.children.forEach(function(ch){const che=ct.querySelector(`[data-bid="${ch.id}"]`);if(che)ch.content=che.textContent||'';});
if(b.type==='columns'&&b.children)b.children.forEach(function(ch){const che=ct.querySelector(`[data-bid="${ch.id}"]`);if(che)ch.content=che.textContent||'';});
const el=ct.querySelector(`[data-bid="${b.id}"]`);if(el&&!el.classList.contains('toggle-title'))b.content=gt(el)||'';else if(el)b.content=gt(el)||'';
if(b.type==='toggle'&&b.children)b.children.forEach(function(ch){const che=ct.querySelector(`[data-bid="${ch.id}"]`);if(che)ch.content=gt(che)||'';});
if(b.type==='columns'&&b.children)b.children.forEach(function(ch){const che=ct.querySelector(`[data-bid="${ch.id}"]`);if(che)ch.content=gt(che)||'';});
}},
render(){const ct=document.getElementById('_blocksCt');if(!ct)return;let html='';for(let i=0;i<this.blocks.length;i++)html+=renderBlock(this.blocks[i],i);ct.innerHTML=html;
// Render math blocks with KaTeX
@@ -1120,7 +1130,8 @@ applyAIBlocks(text){
if(e.key==='Enter'){
if(e.shiftKey){this.dirty=true;this.autoSave();return;}
e.preventDefault();
const text=el.textContent||'',pos=cp(el),before=text.substring(0,pos),after=text.substring(pos);
const sp=splitCaret(el);
const before=sp?sp.before:gt(el)||'',after=sp?sp.after:'';
if(!before.trim()&&!after.trim()){
const nt=(block.type==='bulleted_list'||block.type==='numbered_list'||block.type==='to_do')?block.type:'paragraph';
this.sync();this.blocks.splice(idx+1,0,this.mkB(nt,''));this.render();
@@ -1132,8 +1143,8 @@ applyAIBlocks(text){
setTimeout(()=>{const ne=this.getEl(this.blocks[idx+1]?.id);if(ne){ne.focus();cs(ne);}},60);
this.dirty=true;this.autoSave();return;
}
if(e.key==='ArrowUp'){if(cp(el)===0&&idx>0){e.preventDefault();const prev=this.getEl(this.blocks[idx-1]?.id);if(prev){prev.focus();ce(prev);}}return;}
if(e.key==='ArrowDown'){if(cp(el)>=(el.textContent||'').length&&idx<this.blocks.length-1){e.preventDefault();const next=this.getEl(this.blocks[idx+1]?.id);if(next){next.focus();cs(next);}}return;}
if(e.key==='ArrowUp'){const _sp=splitCaret(el);if(_sp&&!_sp.before.trim()&&idx>0){e.preventDefault();const prev=this.getEl(this.blocks[idx-1]?.id);if(prev){prev.focus();ce(prev);}}return;}
if(e.key==='ArrowDown'){const _sp=splitCaret(el);if(_sp&&!_sp.after.trim()&&idx<this.blocks.length-1){e.preventDefault();const next=this.getEl(this.blocks[idx+1]?.id);if(next){next.focus();cs(next);}}return;}
if(e.key==='Backspace'){if(!(el.textContent||'').trim()){this.sync();if(block.type!=='paragraph'){e.preventDefault();block.type='paragraph';block.content='';this.render();setTimeout(()=>{const ne=this.getEl(block.id);if(ne)ne.focus();},60);}else if(this.blocks.length>1){e.preventDefault();this.removeBlock(idx);}}return;}
if(e.key===' '){const b2=bid;setTimeout(()=>this.checkMd(b2),15);}
if((e.ctrlKey||e.metaKey)&&!e.altKey){const m={b:'bold',i:'italic',u:'underline'};if(m[e.key]){e.preventDefault();document.execCommand(m[e.key]);this.dirty=true;this.autoSave();}}
+139
View File
@@ -0,0 +1,139 @@
"""FlowDeck — tests for app.services.llm_config.fetch_provider_models.
Verifies the live model-list fetching and the chat-capability validation used
for noisy providers (NVIDIA lists its whole catalog on /v1/models, most of
which answers 404 on /v1/chat/completions).
"""
import asyncio
import json
import os
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import pytest
# Pinned BEFORE any app import: app.config.settings is a process-wide singleton
# frozen at first import, and pytest imports test modules during collection —
# before the `client` fixtures of other test files set these env vars.
os.environ["RATE_LIMIT_ENABLED"] = "false"
os.environ["LLM_PROVIDER"] = "offline"
os.environ["DATABASE_URL"] = "sqlite:///:memory:"
os.environ["APP_SECRET_KEY"] = "test-secret-for-tests"
from app.services.llm_config import _is_likely_chat, fetch_provider_models
# Models returned by the mock {"id": ...} list.
MOCK_CATALOG = [
"nvidia/nemotron-3-super-120b-a12b", # chat-capable
"meta/llama-3.1-70b-instruct", # chat-capable
"nvidia/embed-qa-4", # embeddings only → 404 on chat
"nvidia/rerank-qa-mistral-4b", # reranker only → 404 on chat
"black-forest-labs/flux-1-schnell", # image gen → 404 on chat
]
def _make_server(reject_non_chat: bool = True, accept: set | None = None,
fail_all_chat: bool = False):
"""HTTP server mimicking an OpenAI-compatible /models + /chat/completions.
With ``reject_non_chat`` the chat endpoint returns 404 for any model whose
id contains embed/rerank/flux, exactly like NVIDIA's public API.
With ``fail_all_chat`` every chat call returns 404 (probe failure case).
``accept`` is an optional allow-list; when set, only those models answer 200.
"""
accept_models = accept if accept is not None else set(MOCK_CATALOG[:2])
class Handler(BaseHTTPRequestHandler):
def log_message(self, *args): # silence test noise
pass
def do_GET(self):
if self.path.endswith("/models"):
body = json.dumps({"data": [{"id": m} for m in MOCK_CATALOG]}).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
else:
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
def do_POST(self):
if self.path.endswith("/chat/completions"):
n = int(self.headers.get("Content-Length", 0) or 0)
try:
req = json.loads(self.rfile.read(n) or b"{}")
except ValueError:
req = {}
model = req.get("model", "")
works = model in accept_models
if reject_non_chat and any(x in model for x in ("embed", "rerank", "flux")):
works = False
if fail_all_chat:
works = False
self.send_response(200 if works else 404)
self.send_header("Content-Length", "0")
self.end_headers()
else:
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
threading.Thread(target=server.serve_forever, daemon=True).start()
return server
@pytest.fixture
def mock_server():
srv = _make_server()
yield srv
srv.shutdown()
def test_is_likely_chat_filters_obvious_non_llm():
assert _is_likely_chat("nvidia/nemotron-3-super-120b-a12b")
assert _is_likely_chat("meta/llama-3.1-70b-instruct")
assert not _is_likely_chat("nvidia/embed-qa-4")
assert not _is_likely_chat("snowflake/arctic-embed-l")
assert not _is_likely_chat("nvidia/rerank-qa-mistral-4b")
assert not _is_likely_chat("black-forest-labs/flux-1-schnell")
assert not _is_likely_chat("nvidia/tts")
def test_nvidia_only_chat_models_survive(mock_server):
base = f"http://127.0.0.1:{mock_server.server_port}/v1"
models = asyncio.run(fetch_provider_models(
"nvidia", api_key="nvapi-test", api_base=base, timeout=2)
)
assert models == ["nvidia/nemotron-3-super-120b-a12b", "meta/llama-3.1-70b-instruct"]
def test_validation_falls_back_to_name_filter_when_all_fail():
# Everything answers 404 → validated list empty → keep the name-filtered
# candidates instead of wiping the selector.
srv = _make_server(fail_all_chat=True)
try:
base = f"http://127.0.0.1:{srv.server_port}/v1"
models = asyncio.run(fetch_provider_models(
"nvidia", api_key="nvapi-test", api_base=base, timeout=2,
))
finally:
srv.shutdown()
assert set(models) == {"nvidia/nemotron-3-super-120b-a12b", "meta/llama-3.1-70b-instruct"}
def test_other_providers_keep_full_list(mock_server):
# Provider not in _CHAT_VALIDATED_PROVIDERS (openai) is returned verbatim:
# no chat probes are sent (server returns 404 on POST, would break otherwise).
base = f"http://127.0.0.1:{mock_server.server_port}/v1"
models = asyncio.run(fetch_provider_models(
"openai", api_key="sk-test", api_base=base, timeout=2)
)
assert models == MOCK_CATALOG
def test_offline_returns_empty_list():
assert asyncio.run(fetch_provider_models("offline", api_base="", timeout=2)) == []