feat(youtube): InnerTube-first search, transcripts and watch-next related (Steps 15-18)

- InnerTube layer via pinned youtubei.js 18.1.0 (no quota, no key):
  search with merged continuations (unlimited pages), watch-next
  related with LockupView mapping, caption-track discovery
- 3-layer dispatcher (YT_SEARCH_MODE, default innertube-first):
  innertube -> yt-dlp scrape -> official API, graceful errors.yt
- Robust yt-dlp binary resolution (YT_DLP_PATH > PATH > bundled)
  with systematic API fallback (fixes spawn ENOENT in UI)
- Transcript: InnerTube caption discovery (YT_TRANSCRIPT_SOURCE),
  reusing pickTrack/orderedTracks/parseTrackText; yt-dlp fallback kept
- Watch: sidebar uses real watch-next related[] (/api/details),
  title-search fallback for other providers
- Cache: memory LRU + SQLite (youtube_search_cache, youtube_metrics),
  never persist empty pages; /healthz observability; /api/trending
- Includes pending Step 15/16 leftovers in same files (suggest,
  test scripts); unrelated provider adapters left uncommitted
This commit is contained in:
2026-09-25 19:38:46 -04:00
parent 300e5c45b1
commit 4bdcd393fe
25 changed files with 4112 additions and 168 deletions
+21 -1
View File
@@ -1,8 +1,28 @@
# Configuration des clés API pour les adaptateurs de recherche
# Copiez ce fichier en .env et remplissez les valeurs
# YouTube API Key (obligatoire pour la recherche YouTube)
# --- YouTube hybride Step 17-18 : InnerTube (sans clé) > scrape yt-dlp > API officielle ---
# Mode : innertube-first (défaut, 0 quota, pagination illimitée) | innertube-only
# | scrape-first | api-first | scrape-only | api-only
YT_SEARCH_MODE=innertube-first
# Localisation InnerTube (gl=region, hl=langue)
YT_INNERTUBE_GL=FR
YT_INNERTUBE_HL=fr
# Découverte des pistes de sous-titres YouTube : innertube-first (défaut,
# sans spawn yt-dlp) | ytdlp-only | innertube-only
YT_TRANSCRIPT_SOURCE=innertube-first
# Cache scrape (30 min) + binaire yt-dlp (frais requis contre le bot-check)
YT_SCRAPE_TTL_MS=1800000
YT_DLP_TIMEOUT_MS=20000
# YT_DLP_PATH=/usr/local/bin/yt-dlp
# Anti-ban : cookies exportés (volume :ro en Docker) + PO-Token + proxy sortant
# YT_COOKIES_FILE=/cookies/youtube.txt
# YT_PO_TOKEN=
# YT_EGRESS_PROXY=http://proxy:8080
# YouTube API Key (fallback uniquement en scrape-first, obligatoire en api-only)
# Obtenez une clé ici : https://console.developers.google.com/
# Accepte YOUTUBE_API_KEYS en CSV ('k1,k2') ou JSON ('["k1","k2"]') avec rotation auto.
YOUTUBE_API_KEY=your_youtube_api_key_here
# Twitch Client ID (optionnel, pour une recherche Twitch plus complète)
+12
View File
@@ -42,3 +42,15 @@ jobs:
- name: E2E scenarios (search UX)
run: npm run test:search-e2e
- name: Suggest typeahead (Step 15)
run: npm run test:suggest
- name: Transcripts (Step 16)
run: npm run test:transcript
- name: YouTube scrape-first (Step 17, offline)
run: npm run test:ytscrape
- name: YouTube InnerTube (Step 18, offline)
run: npm run test:ytinnertube
+12 -3
View File
@@ -56,20 +56,28 @@ Un seul champ, tous les fournisseurs — avec filtres, raccourcis et deep-links
* **Deep-links** : `/#/search?q=…&providers=yt,ru` relance la recherche filtrée — partageable
* **Fallback préférence** : URL sans `providers` → préférence `defaultProviders` de l’utilisateur → provider actif
* **Accessibilité** : focus trap dans les modals, Esc pour fermer, aria-combobox sur le champ, focus restauré à la fermeture
* **Typeahead requête** : sous l’input, suggestions de requêtes (recherches récentes 🕘 + groupes par provider `YT/DM/…`, sous-chaîne surlignée) — `GET /api/search/suggest?q=…&providers=…&limit=…` (min 2 caractères, debounce 250 ms, cache 5 min, dégradation `[]` par provider) ; ↑/↓/Enter/Tab/Esc, priorité au popover `@`
<!-- TODO: add docs/search-ux.gif (capture des chips, du picker Ctrl+K et du deep-link) -->
### Endpoints API concernés
* `GET /api/search?q=…&providers=yt,dm` — fan-out parallèle, réponse groupée par provider
* `GET /api/search?q=…&providers=yt,dm` — fan-out parallèle, réponse groupée par provider (`page`/`pageSize`/`sort` ; YT sans quota via InnerTube + continuations)
* `GET /api/search/suggest?q=…&providers=yt,dm&limit=10` — typeahead `{ q, groups: { yt: string[], … } }` (cache 5 min, rate-limit)
* `GET /api/details/youtube/:videoId` — métadonnées + `related[]` (watch-next InnerTube, `?related=0` pour désactiver)
* `GET /api/trending?provider=yt&limit=…` — tendances YT sans clé
* `GET /healthz` (alias `/api/healthz`) — mode YT, binaire yt-dlp `binOk`, cache, métriques quota/jour, clés
* `GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=` — transcript `{ lang, available, languages, lines: [{ t, dur, text }] }` (cache 24 h, rate-limit 10/min ; absent → 200 `{ available: false }`, échec → 502 ; YouTube : découverte des pistes via InnerTube, `YT_TRANSCRIPT_SOURCE`)
* `GET/PATCH /api/user/preferences` — `defaultProviders` (tableau JSON, sanitizé serveur)
* `POST /api/telemetry/events` — événements UX anonymes (whitelist : `search_submit`, `provider_picker_open`, `provider_apply`, `at_autocomplete_use`, `quick_menu_open`)
* `POST /api/telemetry/events` — événements UX anonymes (whitelist : `search_submit`, `provider_picker_open`, `provider_apply`, `at_autocomplete_use`, `quick_menu_open`, `suggest_shown`, `suggest_used`)
### Tests
```bash
npm run test:search # unitaires SearchService + parsing @ + picker
npm run test:search-e2e # scénarios e2e (serveur réel isolé)
npm run test:suggest # typeahead : parsing/dédup + contrat /api/search/suggest
npm run test:transcript # transcripts : parseurs json3/vtt + contrat /api/transcript
npm run test:preferences # persistance defaultProviders
npm run test:telemetry # télémétrie minimale
```
@@ -242,7 +250,8 @@ Ajoutez au besoin `log-driver`, `log-opts`, `default-address-pools`, etc.
* ✅ **Téléchargements** intégrés — file d'attente **persistée en SQLite** (survit aux redémarrages API), jobs **par utilisateur** (répertoires isolés, ownership sur status/fichier/cancel), **quota de stockage** configurable (`DOWNLOAD_STORAGE_QUOTA_BYTES`, fenêtre `DOWNLOAD_QUOTA_WINDOW_MS`), **reprise** des jobs échoués/interrrompus (bouton Réessayer), **page Bibliothèque > Téléchargements** (filtres par état, progression live, quota), nettoyage auto des fichiers orphelins au boot
* 🔧 Variables : `DOWNLOAD_MAX_CONCURRENT` (2), `DOWNLOAD_STORAGE_QUOTA_BYTES` (5 GiB), `DOWNLOAD_QUOTA_WINDOW_MS` (30 j), `DOWNLOAD_PROVIDERS` (peertube,odysee)
* ⏳ **Import/Export** playlists (JSON / OPML-like)
* ⏳ **Sous-titres & transcripts** (si dispo API; fallback parsing)
* ✅ **Sous-titres & transcripts** — `GET /api/transcript/:provider/:videoId` (yt-dlp `subtitles`/`automatic_captions`, parsing `json3`/`vtt`, cache 24 h, rate-limit), panneau **Transcript** sur la page Watch (sélecteur de langue, dégradation propre si indisponible)
* ✅ **YouTube sans quota (InnerTube façon SmartTube)** — `youtubei.js` pinné, chaîne `innertube → scrape yt-dlp → API officielle` (`YT_SEARCH_MODE`, défaut `innertube-first`), **pagination illimitée** via continuations (scroll infini, `page=2,3…`), **vidéos connexes** watch-next dans `GET /api/details/youtube/:videoId` → sidebar Watch, `GET /api/trending?provider=yt`, cache mémoire + SQLite, anti-ban (`YT_COOKIES_FILE`, `YT_PO_TOKEN`, `YT_EGRESS_PROXY`), observabilité `/healthz` (mode, binaire `binOk`, cache, quota jour, clés)
* ⏳ **PWA** (installable, offline cache des métadonnées)
* ⏳ **Chromecast / AirPlay**
* ⏳ Mode “TV”
@@ -0,0 +1,18 @@
-- Step 17 : cache persistant recherche YouTube (scrape + API) + métriques quota journalières.
CREATE TABLE IF NOT EXISTS youtube_search_cache (
q_hash TEXT PRIMARY KEY,
q TEXT NOT NULL,
payload_json TEXT NOT NULL,
source TEXT NOT NULL DEFAULT 'scrape',
created_at INTEGER NOT NULL,
expires_at INTEGER NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_yt_cache_exp ON youtube_search_cache(expires_at);
CREATE TABLE IF NOT EXISTS youtube_metrics (
day TEXT PRIMARY KEY,
scrape_calls INTEGER NOT NULL DEFAULT 0,
api_calls INTEGER NOT NULL DEFAULT 0,
quota_units INTEGER NOT NULL DEFAULT 0,
updated_at TEXT NOT NULL
);
+17
View File
@@ -15,6 +15,23 @@ TWITCH_CLIENT_SECRET=votre_client_secret_twitch_ici
# Configuration du cache (en millisecondes)
YT_CACHE_TTL_MS=3600000 # 1 heure par défaut
# --- YouTube hybride : InnerTube (sans clé) > scrape yt-dlp > API officielle ---
# Mode : innertube-first (défaut, 0 quota, pagination illimitée) | innertube-only
# | scrape-first | api-first | scrape-only | api-only
YT_SEARCH_MODE=innertube-first
YT_INNERTUBE_GL=FR
YT_INNERTUBE_HL=fr
# Sous-titres YouTube : innertube-first (défaut) | ytdlp-only | innertube-only
YT_TRANSCRIPT_SOURCE=innertube-first
# Cache scrape/pagination (30 min) + timeout yt-dlp
YT_SCRAPE_TTL_MS=1800000
YT_DLP_TIMEOUT_MS=20000
# Anti-ban (optionnel) : binaire explicite, cookies exportés, PO-Token, proxy sortant
# YT_DLP_PATH=/usr/local/bin/yt-dlp
# YT_COOKIES_FILE=/cookies/youtube.txt
# YT_PO_TOKEN=
# YT_EGRESS_PROXY=http://proxy:8080
# Téléchargements : providers autorisés (liste CSV) + concurrence max par user
DOWNLOAD_PROVIDERS=youtube,dailymotion,twitch,peertube,odysee,rumble
DOWNLOAD_MAX_CONCURRENT=2
+769
View File
@@ -0,0 +1,769 @@
# Guide d'architecture et de développement — Transcripts multi-fournisseurs pour NewTube
**Version :** 1.0
**Date :** 2026-09-25
**Statut :** Proposition d'implémentation
**Périmètre :** Ajout de la fonctionnalité « Transcripts / Sous-titres » sur NewTube, pour les 6 fournisseurs supportés.
---
## Table des matières
1. [Contexte](#1-contexte)
2. [Objectifs et non-objectifs](#2-objectifs-et-non-objectifs)
3. [Architecture globale](#3-architecture-globale)
4. [Backend](#4-backend)
5. [Frontend](#5-frontend)
6. [Référence API](#6-référence-api)
7. [Tests](#7-tests)
8. [Développement pas à pas](#8-développement-pas-à-pas)
9. [Déploiement et configuration](#9-déploiement-et-configuration)
10. [Risques et mitigations](#10-risques-et-mitigations)
11. [Alternatives écartées](#11-alternatives-écartées)
12. [Roadmap](#12-roadmap)
13. [Annexes](#13-annexes)
---
## 1. Contexte
NewTube est un agrégateur multi-fournisseurs qui s'appuie déjà sur `yt-dlp` via `youtube-dl-exec` pour extraire les métadonnées et les formats de téléchargement.
Fournisseurs supportés :
- YouTube
- Dailymotion
- Twitch
- PeerTube
- Odysee
- Rumble
L'analyse du code existant montre que :
- `server/index.mjs:10` importe `youtube-dl-exec`.
- `providerUrlFrom()` (`server/index.mjs:541`) construit une URL normalisée pour les 6 fournisseurs.
- Les routes `/api/details/:provider/:videoId` (`l.988`) et `/api/download/.../formats` (`l.1061`) utilisent déjà `youtubedl(url, { dumpSingleJson, skipDownload })`.
- `yt-dlp` renvoie dans ce même JSON les champs `subtitles` et `automatic_captions`.
- Côté UI, `watch.component.html:154` et `:165` fournissent le motif du panneau « Download » à copier pour un panneau « Transcript ».
- Le README contient déjà une ligne roadmap : `⏳ Sous-titres & transcripts` (`ligne 245`).
La fonctionnalité peut donc être ajoutée **sans nouveau registre de providers**, **sans adaptateur par fournisseur**, et **sans migration de base de données**.
---
## 2. Objectifs et non-objectifs
### 2.1 Objectifs
- Extraire les sous-titres manuels et automatiques via `yt-dlp`.
- Supporter plusieurs langues avec un sélecteur.
- Offrir une dégradation propre quand aucune piste n'existe.
- Mettre en cache les transcripts (immuables) pour éviter les appels répétés.
- Limiter le débit pour éviter les erreurs HTTP 429 de YouTube.
- Ajouter une UI simple dans la page « Watch ».
- Rester compatible avec les 6 fournisseurs sans code spécifique par plateforme.
### 2.2 Non-objectifs
- Traduction automatique des transcripts.
- Transcription audio via LLM ou service tiers.
- Téléchargement de la vidéo complète.
- Authentification OAuth.
- Seek avancé dans la vidéo depuis le transcript (optionnel, phase 2).
- Support garanti des sous-titres sur Twitch, Odysee et Rumble.
---
## 3. Architecture globale
### 3.1 Schéma
```mermaid
flowchart LR
A[Client Watch] -->|GET /api/transcript/:provider/:videoId| B[API NewTube]
B --> C{Cache ?}
C -->|Oui| D[Retour JSON]
C -->|Non| E[providerUrlFrom]
E --> F[yt-dlp dumpSingleJson]
F --> G[subtitles / automatic_captions]
G --> H[pickTrack]
H --> I[Fetch piste json3/vtt]
I --> J[parseJson3 / parseVtt]
J --> K[Mise en cache]
K --> D
D --> A
```
### 3.2 Composants
| Composant | Fichier | Rôle |
|---|---|---|
| Module transcript | `server/transcript.mjs` | Fonctions pures : sélection de piste, parsing json3/vtt |
| Route API | `server/index.mjs` | Endpoint `/api/transcript/...`, cache, rate-limit |
| UI Watch | `watch.component.ts/html` | Bouton, panneau, sélecteur de langue, affichage |
| Tests | `server/tests/transcript.test.mjs` | Tests unitaires et d'intégration |
| Script npm | `package.json` | `test:transcript` |
| Documentation | `README.md` | Case roadmap cochée + bump de version |
### 3.3 Principe directeur
> **Rung 2 : réutiliser, pas réinventer.**
Une seule route, un seul module de parsing, une seule UI. La capacité dépend de la plateforme, mais le code ne change pas.
---
## 4. Backend
### 4.1 Module `server/transcript.mjs`
Fonctions pures, testables sans réseau ni base de données.
#### 4.1.1 `pickTrack(json, lang)`
Sélectionne la meilleure piste disponible.
Priorités :
1. Sous-titres manuels (`subtitles`) dans la langue demandée.
2. Sous-titres manuels dans une langue proche (`fr-*`).
3. Sous-titres automatiques (`automatic_captions`) dans la langue demandée.
4. Sous-titres automatiques dans une langue proche.
5. Première piste disponible.
```js
export function pickTrack(json, lang = 'fr') {
const manual = json.subtitles || {};
const auto = json.automatic_captions || {};
const all = { ...auto, ...manual };
const languages = Object.keys(all);
if (!languages.length) {
return { track: null, languages: [], lang: null };
}
const findLang = (dict, code) =>
dict[code] || Object.keys(dict).find(k => k.startsWith(code + '-'));
const chosenLang =
findLang(manual, lang) ? lang :
findLang(auto, lang) ? lang :
languages[0];
const track =
manual[chosenLang] ||
auto[chosenLang] ||
manual[languages[0]] ||
auto[languages[0]];
return { track, languages, lang: chosenLang };
}
```
#### 4.1.2 `parseJson3(data)`
Convertit le format `json3` de YouTube en lignes normalisées.
```js
export function parseJson3(data) {
const events = data.events || [];
return events
.filter(e => e.segs)
.map(e => ({
t: (e.tStartMs || 0) / 1000,
dur: (e.dDurationMs || 0) / 1000,
text: e.segs.map(s => s.utf8).join('').trim()
}))
.filter(line => line.text);
}
```
#### 4.1.3 `parseVtt(text)`
Convertit un fichier VTT en lignes normalisées.
```js
export function parseVtt(text) {
const lines = [];
const blocks = text.split(/\n\n+/);
for (const block of blocks) {
const match = block.match(/(\d{2}):(\d{2}):(\d{2})\.(\d{3}) --> (\d{2}):(\d{2}):(\d{2})\.(\d{3})/);
if (!match) continue;
const start =
parseInt(match[1]) * 3600 +
parseInt(match[2]) * 60 +
parseInt(match[3]) +
parseInt(match[4]) / 1000;
const end =
parseInt(match[5]) * 3600 +
parseInt(match[6]) * 60 +
parseInt(match[7]) +
parseInt(match[8]) / 1000;
const content = block
.split('\n')
.filter(line => !line.includes('-->') && !/^\d+$/.test(line))
.join(' ')
.trim();
if (content) {
lines.push({ t: start, dur: end - start, text: content });
}
}
return lines;
}
```
#### 4.1.4 Structure de retour normalisée
```json
{
"lang": "fr",
"available": true,
"languages": ["fr", "en", "es"],
"lines": [
{ "t": 0, "dur": 2.5, "text": "Bonjour à tous" },
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue dans cette vidéo" }
]
}
```
En l'absence de piste :
```json
{
"lang": null,
"available": false,
"languages": [],
"lines": []
}
```
---
### 4.2 Route API dans `server/index.mjs`
#### 4.2.1 Signature
```http
GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=
```
#### 4.2.2 Flux d'exécution
1. Valider le provider.
2. Construire la clé de cache : `transcript:${provider}:${videoId}:${lang}`.
3. Vérifier le cache.
4. Si absent :
- `providerUrlFrom()` construit l'URL.
- `youtubedl(url, { dumpSingleJson: true, skipDownload: true })`.
- `pickTrack(json, lang)`.
- Si aucune piste : répondre `{ available: false }` avec HTTP 200.
- Sinon : `fetch(track.url)`.
- Parser selon le format (`json3` ou `vtt`).
- Mettre en cache.
5. Répondre.
#### 4.2.3 Pseudo-code
```js
app.get('/api/transcript/:provider/:videoId', channelsLimiter, async (req, res) => {
const { provider, videoId } = req.params;
const { lang = 'fr', instance, slug, sourceUrl } = req.query;
const cacheKey = `transcript:${provider}:${videoId}:${lang}`;
const cached = transcriptCache.get(cacheKey);
if (cached) return res.json(cached);
try {
const url = providerUrlFrom(provider, videoId, { instance, slug, sourceUrl });
const json = await youtubedl(url, {
dumpSingleJson: true,
skipDownload: true,
});
const { track, languages, lang: chosenLang } = pickTrack(json, lang);
if (!track) {
const empty = { lang: null, available: false, languages: [], lines: [] };
transcriptCache.set(cacheKey, empty, TTL_LONG);
return res.json(empty);
}
const response = await fetch(track.url);
const text = await response.text();
const lines = track.ext === 'json3'
? parseJson3(JSON.parse(text))
: parseVtt(text);
const result = {
lang: chosenLang,
available: true,
languages,
lines,
};
transcriptCache.set(cacheKey, result, TTL_LONG);
res.json(result);
} catch (err) {
console.error('Transcript error:', err);
res.status(502).json({ available: false, error: 'transcript_fetch_failed' });
}
});
```
---
### 4.3 Cache
- **Type :** `Map` avec TTL, comme `ytCache` (`server/index.mjs:838`).
- **Clé :** `transcript:${provider}:${videoId}:${lang}`.
- **TTL :** long (ex. 24 h) car les transcripts sont immuables.
- **Évolution :** pour un déploiement multi-instances, remplacer par Redis ou Memcached.
```js
const transcriptCache = {
data: new Map(),
get(key) {
const entry = this.data.get(key);
if (!entry) return null;
if (Date.now() > entry.expires) {
this.data.delete(key);
return null;
}
return entry.value;
},
set(key, value, ttlMs) {
this.data.set(key, { value, expires: Date.now() + ttlMs });
}
};
```
---
### 4.4 Rate limiting
- Réutiliser le motif `channelsLimiter`.
- Limiter par IP, par exemple 10 requêtes / minute sur `/api/transcript/...`.
- En cas de HTTP 429 de YouTube :
- Ne pas réessayer immédiatement.
- Attendre 20 s minimum.
- Utiliser un backoff exponentiel.
- Si le problème persiste, activer le fallback.
---
### 4.5 Fallback en cas de 429 persistant
Si le endpoint `timedtext` de YouTube renvoie trop de 429 :
1. Laisser `yt-dlp` écrire le VTT dans un dossier temporaire.
2. Lire le fichier.
3. Parser avec `parseVtt()`.
Pattern déjà utilisé pour les downloads : `youtubedl.exec`.
```js
const tmp = path.join(os.tmpdir(), `transcript-${Date.now()}.vtt`);
await youtubedl.exec(url, {
writeAutoSub: true,
subLang: lang,
subFormat: 'vtt',
output: tmp,
});
const text = await fs.readFile(tmp, 'utf8');
const lines = parseVtt(text);
```
---
## 5. Frontend
### 5.1 UI dans `watch.component`
#### 5.1.1 Bouton
Ajouter un bouton « Transcript » à côté du bouton « Download ».
```html
<button class="btn-transcript" (click)="toggleTranscript()">
Transcript
</button>
```
#### 5.1.2 Panneau
Copier le motif du panneau Download (`watch.component.html:165`).
```html
<div *ngIf="transcriptOpen" class="transcript-panel">
<div class="transcript-header">
<h3>Transcript</h3>
<select [(ngModel)]="selectedLang" (change)="loadTranscript()">
<option *ngFor="let lang of transcriptLanguages" [value]="lang">
{{ lang }}
</option>
</select>
</div>
<div *ngIf="transcriptLoading" class="transcript-loading">
Chargement…
</div>
<div *ngIf="!transcriptLoading && !transcriptAvailable" class="transcript-empty">
Aucun sous-titre disponible pour cette vidéo.
</div>
<div *ngIf="transcriptAvailable" class="transcript-lines">
<div *ngFor="let line of transcriptLines" class="transcript-line">
<span class="transcript-time">{{ line.t | duration }}</span>
<span class="transcript-text">{{ line.text }}</span>
</div>
</div>
</div>
```
#### 5.1.3 Composant TypeScript
```ts
transcriptOpen = false;
transcriptLoading = false;
transcriptAvailable = false;
transcriptLanguages: string[] = [];
transcriptLines: { t: number; dur: number; text: string }[] = [];
selectedLang = 'fr';
toggleTranscript() {
this.transcriptOpen = !this.transcriptOpen;
if (this.transcriptOpen && !this.transcriptLines.length) {
this.loadTranscript();
}
}
async loadTranscript() {
this.transcriptLoading = true;
const res = await fetch(
`/api/transcript/${this.provider}/${this.videoId}?lang=${this.selectedLang}`
);
const data = await res.json();
this.transcriptAvailable = data.available;
this.transcriptLanguages = data.languages || [];
this.transcriptLines = data.lines || [];
this.transcriptLoading = false;
}
```
#### 5.1.4 Gestion de l'absence de sous-titres
- Si `available: false`, afficher un message clair.
- Le bouton peut rester visible mais le panneau indique l'absence.
- Ne jamais casser la page Watch.
---
### 5.2 Variante UI #1 — recommandée
- Affichage du transcript.
- Sélecteur de langue.
- Pas de seek.
**Avantages :** simple, robuste, peu de code.
---
### 5.3 Variante UI #2 — optionnelle
- Comme #1, plus : clic sur une ligne = seek dans la vidéo.
**Implémentation :**
- Player natif : `seekBy` dans `video-player.component.ts:112`.
- Iframe YouTube : `seekTo` dans `iframe-progress.service.ts:104`.
**Inconvénient :** le seek iframe n'est attaché que sous certaines conditions. Il faut brancher par provider. Sensiblement plus de code.
**Recommandation :** garder #2 pour une phase ultérieure.
---
## 6. Référence API
### 6.1 Endpoint
```http
GET /api/transcript/:provider/:videoId
```
### 6.2 Paramètres
| Nom | Type | Requis | Description |
|---|---|---|---|
| `provider` | string | Oui | `youtube`, `dailymotion`, `twitch`, `peertube`, `odysee`, `rumble` |
| `videoId` | string | Oui | Identifiant de la vidéo |
| `lang` | string | Non | Langue souhaitée (ex. `fr`, `en`). Défaut : `fr` |
| `instance` | string | Non | Pour PeerTube |
| `slug` | string | Non | Pour PeerTube |
| `sourceUrl` | string | Non | Pour Odysee / Rumble |
### 6.3 Réponse succès
```json
{
"lang": "fr",
"available": true,
"languages": ["fr", "en"],
"lines": [
{ "t": 0, "dur": 2.5, "text": "Bonjour" },
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue" }
]
}
```
### 6.4 Réponse absence de sous-titres
```json
{
"lang": null,
"available": false,
"languages": [],
"lines": []
}
```
HTTP 200.
### 6.5 Réponse erreur
```json
{
"available": false,
"error": "transcript_fetch_failed"
}
```
HTTP 502.
---
## 7. Tests
### 7.1 Tests unitaires
Fichier : `server/tests/transcript.test.mjs`
- `pickTrack()` :
- manuel prioritaire sur auto.
- `fr` → `fr-*` → autre langue.
- aucun track → `available: false`.
- `parseJson3()` :
- events valides.
- events sans `segs`.
- texte vide filtré.
- `parseVtt()` :
- blocs valides.
- timestamps corrects.
- contenu multi-lignes.
### 7.2 Tests d'intégration
- Mocker `youtubedl` pour renvoyer un JSON avec `subtitles`.
- Vérifier la route `/api/transcript/...`.
- Vérifier le cache : deuxième appel ne déclenche pas `youtubedl`.
- Vérifier le rate-limit.
### 7.3 Tests manuels
| Provider | Vidéo testée | Résultat attendu |
|---|---|---|
| YouTube | Vidéo avec auto-captions | `available: true`, ~100 langues |
| Dailymotion | Vidéo avec pistes | `available: true` |
| PeerTube | Vidéo framatube | souvent `available: false` |
| Twitch | VOD | `available: false` |
| Odysee | Vidéo | `available: false` |
| Rumble | Vidéo | `available: false` |
### 7.4 Script npm
```json
{
"scripts": {
"test:transcript": "node --test server/tests/transcript.test.mjs"
}
}
```
---
## 8. Développement pas à pas
### Checklist
- [ ] Créer `server/transcript.mjs` avec `pickTrack`, `parseJson3`, `parseVtt`.
- [ ] Ajouter la route `GET /api/transcript/:provider/:videoId` dans `server/index.mjs`.
- [ ] Ajouter le cache `transcriptCache` avec TTL long.
- [ ] Ajouter le rate-limit sur le motif `channelsLimiter`.
- [ ] Ajouter la gestion des erreurs 429 et le fallback temp dir.
- [ ] Créer `server/tests/transcript.test.mjs`.
- [ ] Ajouter `test:transcript` dans `package.json`.
- [ ] Ajouter le bouton et le panneau dans `watch.component.html`.
- [ ] Ajouter la logique dans `watch.component.ts`.
- [ ] Mettre à jour le README : cocher `⏳ Sous-titres & transcripts`.
- [ ] Bumper la version selon la règle du projet.
- [ ] Tester manuellement sur YouTube et Dailymotion.
- [ ] Vérifier la dégradation propre sur Twitch, Odysee, Rumble.
**Volume estimé :** 150 à 200 lignes, 4 fichiers.
---
## 9. Déploiement et configuration
### 9.1 Variables d'environnement
| Variable | Description | Défaut |
|---|---|---|
| `YTDLP_PATH` | Chemin vers `yt-dlp` | `yt-dlp` |
| `TRANSCRIPT_CACHE_TTL` | TTL du cache en ms | `86400000` (24 h) |
| `TRANSCRIPT_RATE_LIMIT` | Requêtes / minute / IP | `10` |
| `TRANSCRIPT_FALLBACK_TMP` | Activer le fallback temp dir | `false` |
### 9.2 Mise à jour de `yt-dlp`
- `yt-dlp` évolue vite.
- Prévoir une mise à jour régulière.
- Surveiller les breaking changes.
### 9.3 Monitoring
- Compter les HTTP 429 sur `timedtext`.
- Compter les `available: false` par provider.
- Mesurer le temps de réponse de `/api/transcript/...`.
- Alerter si le taux d'échec dépasse un seuil.
---
## 10. Risques et mitigations
| Risque | Impact | Mitigation |
|---|---|---|
| ToS des plateformes | Élevé | Usage personnel / auto-hébergé, ne pas revendre |
| HTTP 429 YouTube | Moyen | Cache long, rate-limit, fallback temp dir |
| `yt-dlp` cassé | Élevé | Mise à jour régulière, tests de non-régression |
| Providers sans sous-titres | Faible | Dégradation propre `available: false` |
| Cache mémoire non partagé | Moyen | Redis en production multi-instances |
| Seek iframe complexe | Faible | Reporter en phase 2 |
---
## 11. Alternatives écartées
| Alternative | Raison du rejet |
|---|---|
| YouTube Data API `captions.download` | Nécessite OAuth du propriétaire |
| Un implémentation par provider | 6 chemins de code pour le même résultat |
| Fetch navigateur `timedtext` | CORS non garanti + casse le proxy clé API |
| Services tiers / LLM de transcription | Dépendance, coût, vie privée, YAGNI |
| Scrapers custom par plateforme | Maintenance impossible |
---
## 12. Roadmap
### Phase 1 — Base (recommandée)
- Route `/api/transcript/...`
- Module `transcript.mjs`
- UI #1 : affichage + sélecteur de langue
- Cache + rate-limit
- Tests
### Phase 2 — Confort
- UI #2 : clic sur une ligne = seek
- Branchement par provider
- Meilleure gestion des timecodes
### Phase 3 — Scalabilité
- Cache Redis
- Proxies rotatifs
- Monitoring avancé
- Support de nouveaux formats
---
## 13. Annexes
### 13.1 Format `json3`
Exemple simplifié :
```json
{
"events": [
{
"tStartMs": 0,
"dDurationMs": 2500,
"segs": [{ "utf8": "Bonjour" }]
},
{
"tStartMs": 2500,
"dDurationMs": 3100,
"segs": [{ "utf8": "Bienvenue" }]
}
]
}
```
### 13.2 Format VTT
Exemple simplifié :
```vtt
WEBVTT
00:00:00.000 --> 00:00:02.500
Bonjour
00:00:02.500 --> 00:00:05.600
Bienvenue
```
### 13.3 Exemple de réponse complète
```json
{
"lang": "fr",
"available": true,
"languages": ["fr", "en", "es"],
"lines": [
{ "t": 0, "dur": 2.5, "text": "Bonjour à tous" },
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue dans cette vidéo" },
{ "t": 5.6, "dur": 4.2, "text": "Aujourd'hui, nous allons voir..." }
]
}
```
### 13.4 Arborescence des fichiers modifiés
```text
server/
index.mjs # route + cache + rate-limit
transcript.mjs # fonctions pures
tests/
transcript.test.mjs # tests
watch.component.ts # logique UI
watch.component.html # panneau Transcript
package.json # script test:transcript
README.md # roadmap + version
```
---
**Fin du guide.**
+36
View File
@@ -34,6 +34,7 @@
"rxjs": "^7.8.2",
"tailwindcss": "latest",
"youtube-dl-exec": "^3.0.0",
"youtubei.js": "^18.1.0",
"zone.js": "~0.15.1"
},
"devDependencies": {
@@ -965,6 +966,12 @@
"node": ">=6.9.0"
}
},
"node_modules/@bufbuild/protobuf": {
"version": "2.15.0",
"resolved": "https://registry.npmjs.org/@bufbuild/protobuf/-/protobuf-2.15.0.tgz",
"integrity": "sha512-DAheWUkVr/SJTWCc+lg9dhY0eN4SaWlf4+bG1KzHeXbnqt0AfB/NX0Z+VunGlM1ki1B4zVvye27MpKh/svySUA==",
"license": "(Apache-2.0 AND BSD-3-Clause)"
},
"node_modules/@cspotcode/source-map-support": {
"version": "0.8.1",
"resolved": "https://registry.npmjs.org/@cspotcode/source-map-support/-/source-map-support-0.8.1.tgz",
@@ -5672,6 +5679,12 @@
}
}
},
"node_modules/fflate": {
"version": "0.8.3",
"resolved": "https://registry.npmjs.org/fflate/-/fflate-0.8.3.tgz",
"integrity": "sha512-tbZNuJrLwGUp3zshBtdy4W+ORxZuIh8a5ilyIEQDC5rY1f3U20JMry0Ll3WBzU58EZKsEuJFXhb5gwv8CsPvgA==",
"license": "MIT"
},
"node_modules/ffmpeg-static": {
"version": "5.2.0",
"resolved": "https://registry.npmjs.org/ffmpeg-static/-/ffmpeg-static-5.2.0.tgz",
@@ -6943,6 +6956,15 @@
"integrity": "sha512-abv/qOcuPfk3URPfDzmZU1LKmuw8kT+0nIHvKrKgFrwifol/doWcdA4ZqsWQ8ENrFKkd67Mfpo/LovbIUsbt3w==",
"license": "MIT"
},
"node_modules/meriyah": {
"version": "7.3.3",
"resolved": "https://registry.npmjs.org/meriyah/-/meriyah-7.3.3.tgz",
"integrity": "sha512-uE5cnoNj+UYhoMdZDuymCzr5TzEuE3ZF2C4kn6Z76rLhkwnvZ/+6dUZnM41T3fifyuhW0+n8vK7eq6DKJxUfPA==",
"license": "ISC",
"engines": {
"node": ">=20.0.0"
}
},
"node_modules/methods": {
"version": "1.1.2",
"resolved": "https://registry.npmjs.org/methods/-/methods-1.1.2.tgz",
@@ -9992,6 +10014,20 @@
"node": ">= 18"
}
},
"node_modules/youtubei.js": {
"version": "18.1.0",
"resolved": "https://registry.npmjs.org/youtubei.js/-/youtubei.js-18.1.0.tgz",
"integrity": "sha512-4Goo3ZeO6/kE8Mq67eC8Z9K1rijeqRwyKxf/YM4lqf8bB4r/s8JC0V3YFx7+l/AJR3LnAsulTsiIj+AFe0cduA==",
"funding": [
"https://github.com/sponsors/LuanRT"
],
"license": "MIT",
"dependencies": {
"@bufbuild/protobuf": "^2.0.0",
"fflate": "^0.8.2",
"meriyah": "^7.3.1"
}
},
"node_modules/zod": {
"version": "3.25.76",
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
+9 -3
View File
@@ -15,7 +15,12 @@
"test:telemetry": "node ./server/tests/telemetry.test.mjs",
"test:downloads": "node ./server/tests/download_jobs.test.mjs",
"test:subscriptions": "node --loader ts-node/esm --experimental-specifier-resolution=node src/services/subscriptions.service.spec.ts",
"test:search": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search.service.spec.ts && node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search-components.spec.ts"
"test:search": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search.service.spec.ts && node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search-components.spec.ts",
"test:suggest": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/suggest.spec.ts && node ./server/tests/suggest.test.mjs",
"test:transcript": "node --test server/tests/transcript.test.mjs",
"test:ytscrape": "node server/tests/youtube-scrape.test.mjs",
"test:ytinnertube": "node server/tests/youtube-innertube.test.mjs",
"ytdlp:update": "yt-dlp -U || python3 -m yt_dlp -U || echo \"yt-dlp update: installez yt-dlp puis relancez\""
},
"dependencies": {
"@angular/build": "^20.1.0",
@@ -44,12 +49,13 @@
"rxjs": "^7.8.2",
"tailwindcss": "latest",
"youtube-dl-exec": "^3.0.0",
"youtubei.js": "18.1.0",
"zone.js": "~0.15.1"
},
"devDependencies": {
"@types/node": "^22.14.0",
"ts-node": "^10.9.2",
"typescript": "~5.8.2",
"vite": "^6.2.0",
"ts-node": "^10.9.2"
"vite": "^6.2.0"
}
}
+71
View File
@@ -998,3 +998,74 @@ function subscriptionRowToDto(row) {
channel,
};
}
// -------------------- Step 17 : cache persistant recherche YouTube --------------------
function ensureYoutubeCacheTables() {
try {
db.exec(`CREATE TABLE IF NOT EXISTS youtube_search_cache (
q_hash TEXT PRIMARY KEY, q TEXT NOT NULL, payload_json TEXT NOT NULL,
source TEXT NOT NULL DEFAULT 'scrape', created_at INTEGER NOT NULL, expires_at INTEGER NOT NULL
);`);
db.exec(`CREATE INDEX IF NOT EXISTS idx_yt_cache_exp ON youtube_search_cache(expires_at);`);
db.exec(`CREATE TABLE IF NOT EXISTS youtube_metrics (
day TEXT PRIMARY KEY, scrape_calls INTEGER NOT NULL DEFAULT 0,
api_calls INTEGER NOT NULL DEFAULT 0, quota_units INTEGER NOT NULL DEFAULT 0, updated_at TEXT NOT NULL
);`);
} catch {}
}
ensureYoutubeCacheTables();
export function getCachedYoutubeSearch(qHash) {
try {
ensureYoutubeCacheTables();
const row = db.prepare(`SELECT payload_json AS payload, source, expires_at AS exp FROM youtube_search_cache WHERE q_hash = ?`).get(qHash);
if (!row) return null;
if (Date.now() >= Number(row.exp || 0)) {
try { db.prepare(`DELETE FROM youtube_search_cache WHERE q_hash = ?`).run(qHash); } catch {}
return null;
}
try { return { items: JSON.parse(String(row.payload || '[]')), source: row.source }; } catch { return null; }
} catch { return null; }
}
export function setCachedYoutubeSearch(qHash, q, items, source, ttlMs) {
try {
ensureYoutubeCacheTables();
const now = Date.now();
db.prepare(`INSERT INTO youtube_search_cache (q_hash, q, payload_json, source, created_at, expires_at)
VALUES (?, ?, ?, ?, ?, ?)
ON CONFLICT(q_hash) DO UPDATE SET q=excluded.q, payload_json=excluded.payload_json,
source=excluded.source, created_at=excluded.created_at, expires_at=excluded.expires_at`)
.run(qHash, String(q || '').slice(0, 300), JSON.stringify(items || []), source, now, now + ttlMs);
// Cap : garde 2000 entrées les plus fraîches
try { db.exec(`DELETE FROM youtube_search_cache WHERE q_hash NOT IN (SELECT q_hash FROM youtube_search_cache ORDER BY expires_at DESC LIMIT 2000)`); } catch {}
} catch {}
}
export function pruneYoutubeCache() {
try { db.prepare(`DELETE FROM youtube_search_cache WHERE expires_at <= ?`).run(Date.now()); } catch {}
}
export function incYoutubeMetrics({ scrapeCalls = 0, apiCalls = 0, quotaUnits = 0 } = {}) {
try {
ensureYoutubeCacheTables();
const day = new Date().toISOString().slice(0, 10);
db.prepare(`INSERT INTO youtube_metrics (day, scrape_calls, api_calls, quota_units, updated_at)
VALUES (?, ?, ?, ?, ?)
ON CONFLICT(day) DO UPDATE SET scrape_calls = scrape_calls + ?, api_calls = api_calls + ?,
quota_units = quota_units + ?, updated_at = excluded.updated_at`)
.run(day, scrapeCalls, apiCalls, quotaUnits, new Date().toISOString(), scrapeCalls, apiCalls, quotaUnits);
} catch {}
}
export function getYoutubeMetricsToday() {
try {
ensureYoutubeCacheTables();
const day = new Date().toISOString().slice(0, 10);
return db.prepare(`SELECT * FROM youtube_metrics WHERE day = ?`).get(day) || { day, scrape_calls: 0, api_calls: 0, quota_units: 0 };
} catch { return { scrape_calls: 0, api_calls: 0, quota_units: 0 }; }
}
export function countYoutubeCacheRows() {
try { return db.prepare(`SELECT COUNT(1) AS n FROM youtube_search_cache`).get()?.n || 0; } catch { return 0; }
}
+458 -3
View File
@@ -7,12 +7,16 @@ import bcrypt from 'bcryptjs';
import jwt from 'jsonwebtoken';
import fs from 'node:fs';
import path from 'node:path';
import youtubedl from 'youtube-dl-exec';
import youtubedlPkg, { create as createYtDlp } from 'youtube-dl-exec';
import { execFile as execFileCb } from 'node:child_process';
import { promisify } from 'node:util';
import { fileURLToPath as serverFileURLToPath } from 'node:url';
import ffmpegPath from 'ffmpeg-static';
import * as cheerio from 'cheerio';
import axios from 'axios';
import rumbleRouter from './rumble.mjs';
import { providerRegistry, validateProviders } from './providers/registry.mjs';
import { pickTrack, parseTrackText, parseVtt, orderedTracks, normalizeTranscriptProvider, transcriptTrackExt, looksLikeHtmlError } from './transcript.mjs';
import {
getUserByUsername,
getUserById,
@@ -73,9 +77,53 @@ import {
} from './db.mjs';
import { getChannelAdapter, setTwitchTokenProvider } from './providers/channel-registry.mjs';
import { fetchChannelContent } from './providers/channel-content.mjs';
import { getSearchMode, getYtDlpBin, hasCookiesFile, metricsSnapshot } from './providers/youtube-common.mjs';
import { ytScrapeCacheStats } from './providers/youtube.mjs';
const app = express();
const PORT = Number(process.env.PORT || 4000);
// yt-dlp: prefer the newest binary available. The copy bundled with
// youtube-dl-exec goes stale (YouTube then answers "The page needs to be
// reloaded" to dump-single-json), while a system install is usually fresher.
// `YT_DLP_PATH` wins, then `yt-dlp` on PATH, then the bundled binary.
const execFileAsync = promisify(execFileCb);
async function binaryVersion(bin) {
try {
const { stdout } = await execFileAsync(bin, ['--version'], { timeout: 15000 });
return String(stdout || '').trim().split('\n')[0].trim();
} catch {
return null;
}
}
let youtubedl = youtubedlPkg;
let ytDlpInfo = 'bundled';
try {
let bundledBin = null;
try {
const u = new URL('../node_modules/youtube-dl-exec/bin/', import.meta.url);
const cand = path.join(String(serverFileURLToPath(u)), process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp');
if (fs.existsSync(cand)) bundledBin = cand;
} catch {}
const bundledVer = bundledBin ? await binaryVersion(bundledBin) : null;
const candidates = [];
if (process.env.YT_DLP_PATH) candidates.push(process.env.YT_DLP_PATH);
candidates.push(process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp', 'yt-dlp');
let best = null;
for (const cand of candidates) {
if (!cand) continue;
const ver = await binaryVersion(cand);
if (ver && (!best || ver > best.ver)) best = { bin: cand, ver };
}
if (best && (!bundledVer || best.ver >= bundledVer)) {
youtubedl = createYtDlp(best.bin);
ytDlpInfo = `${best.bin} (${best.ver})`;
} else if (bundledVer) {
ytDlpInfo = `bundled (${bundledVer})`;
}
} catch (e) {
console.warn('[config] yt-dlp binary selection failed, using bundled:', e?.message || e);
}
console.log(`[config] yt-dlp: ${ytDlpInfo}`);
const IS_PROD = String(process.env.NODE_ENV || '').toLowerCase() === 'production';
const JWT_SECRET = process.env.JWT_SECRET || 'dev-secret-change-me';
if (!process.env.JWT_SECRET) {
@@ -1006,6 +1054,14 @@ r.get('/peertube/:instance/*', async (req, res) => {
});
// -------------------- Generic video details (GET) --------------------
// Cache mémoire pour les vidéos connexes InnerTube (watch-next) : TTL 1h, LRU 200.
const YT_RELATED_TTL_MS = Number(process.env.YT_RELATED_TTL_MS || 60 * 60 * 1000);
const ytRelatedCache = new Map();
function ytRelatedCacheSet(key, items) {
if (ytRelatedCache.has(key)) ytRelatedCache.delete(key);
ytRelatedCache.set(key, { ts: Date.now(), items });
while (ytRelatedCache.size > 200) { const o = ytRelatedCache.keys().next().value; if (o === undefined) break; ytRelatedCache.delete(o); }
}
// Returns metadata such as title, description, uploader, thumbnail, duration and views for a provider/videoId
// Supports query params similar to download endpoints: instance (PeerTube), slug (Odysee), sourceUrl (direct)
r.get('/details/:provider/:videoId', async (req, res) => {
@@ -1071,6 +1127,22 @@ r.get('/details/:provider/:videoId', async (req, res) => {
url,
type: 'video',
};
// Step 18 : vidéos connexes façon SmartTube (watch-next InnerTube, best-effort, 0 quota).
// N'implique que YouTube ; toute erreur -> `related: []`, la réponse reste 200.
if (String(provider) === 'youtube' && req.query.related !== '0') {
try {
const { getRelatedViaInnerTube } = await import('./providers/youtube-innertube.mjs');
const relKey = `related:${videoId}`;
const cached = ytRelatedCache.get(relKey);
if (cached && (Date.now() - cached.ts) < YT_RELATED_TTL_MS) {
out.related = cached.items;
} else {
const items = await getRelatedViaInnerTube(videoId, 24).catch(() => []);
out.related = items;
ytRelatedCacheSet(relKey, items);
}
} catch { out.related = []; }
}
return res.json(out);
} catch (e) {
return res.status(500).json({ error: 'details_failed', details: String(e?.message || e) });
@@ -1684,7 +1756,7 @@ r.post('/telemetry/events', authMiddleware, telemetryLimiter, (req, res) => {
const { event, meta } = req.body || {};
if (!event || typeof event !== 'string') return res.status(400).json({ error: 'event_required' });
// Whitelist known event names to keep the table clean
const allowed = new Set(['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open']);
const allowed = new Set(['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open', 'suggest_shown', 'suggest_used']);
if (!allowed.has(event)) return res.status(400).json({ error: 'unknown_event' });
const row = insertTelemetryEvent({ userId: req.user.id, event, meta: (meta && typeof meta === 'object') ? meta : null });
return res.status(201).json(row || { ok: true });
@@ -1704,7 +1776,7 @@ r.get('/telemetry/events', authMiddleware, (req, res) => {
r.get('/telemetry/summary', authMiddleware, (req, res) => {
const since = typeof req.query.since === 'string' ? req.query.since : undefined;
const events = {};
for (const name of ['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open']) {
for (const name of ['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open', 'suggest_shown', 'suggest_used']) {
events[name] = countTelemetryEvents({ event: name, since });
}
return res.json({ events });
@@ -1944,6 +2016,53 @@ r.get('/img/odysee', async (req, res) => {
app.use('/api', r);
// Health endpoint for container checks
app.get('/api/health', (_req, res) => res.json({ status: 'ok' }));
// Step 17 : observabilité YouTube (mode, yt-dlp, cache, quota). Aucun secret exposé.
app.get(['/healthz', '/api/healthz'], async (_req, res) => {
try {
const [{ getYoutubeMetricsToday, countYoutubeCacheRows }, common] =
await Promise.all([import('./db.mjs'), import('./providers/youtube-common.mjs')]);
let ytdlpVersion = null;
let resolvedBin = null;
try {
resolvedBin = await common.resolveYtDlpBin();
const { stdout } = await execFileAsync(resolvedBin, ['--version'], { timeout: 10000 });
ytdlpVersion = String(stdout || '').trim().split('\n')[0].trim() || null;
} catch {}
const keys = common.getYouTubeKeys();
const banned = [];
try { for (const [k, until] of ytKeyBans || []) if (Date.now() < until) banned.push(`...${String(k).slice(-4)}`); } catch {}
res.json({
status: 'ok',
youtube: {
mode: getSearchMode(),
ytdlp: { bin: resolvedBin || getYtDlpBin(), version: ytdlpVersion, info: ytDlpInfo, binOk: Boolean(ytdlpVersion) },
antiban: {
cookiesFile: hasCookiesFile(),
poToken: Boolean(String(process.env.YT_PO_TOKEN || '').trim()),
egressProxy: Boolean(String(process.env.YT_EGRESS_PROXY || '').trim()),
},
cache: { ...ytScrapeCacheStats(), sqliteRows: countYoutubeCacheRows() },
metrics: { ...metricsSnapshot(), today: getYoutubeMetricsToday() },
keys: { count: keys.length, banned },
},
});
} catch (e) {
res.status(500).json({ status: 'error', error: String(e?.message || e) });
}
});
// Step 17 : trending YouTube sans clé (scrape) avec fallback [] propre.
app.get('/api/trending', async (req, res) => {
try {
const provider = String(req.query.provider || 'yt');
const limit = Math.min(50, Math.max(1, Number(req.query.limit || 24)));
if (provider !== 'yt') return res.status(400).json({ error: 'only yt supported in phase 1' });
const { getTrendingViaScrape } = await import('./providers/youtube-scrape.mjs');
const items = await getTrendingViaScrape(limit);
return res.json({ provider, items });
} catch (e) {
return res.json({ provider: 'yt', items: [], error: String(e?.code || e?.message || 'trending_failed') });
}
});
// Alias to support Angular dev proxy paths in both dev and production builds
app.use('/proxy/api', r);
// Mount dedicated Rumble router (browse, search, video)
@@ -2170,6 +2289,342 @@ app.get('/api/search', async (req, res) => {
}
});
// -------------------- Query typeahead suggestions (Step 15) --------------------
// GET /api/search/suggest?q=…&providers=yt,dm&limit=10 -> { q, groups: { yt: string[], dm: string[] } }
// Fan-out over provider `suggest()` handlers; providers without one degrade to [] (never 500).
const SUGGEST_CACHE_TTL_MS = Number(process.env.SUGGEST_CACHE_TTL_MS || 5 * 60 * 1000);
const SUGGEST_CACHE_MAX_ENTRIES = 500;
/** @type {Map<string, { ts: number, data: any }>} */
const suggestCache = new Map();
function suggestCacheGet(key) {
const hit = suggestCache.get(key);
if (!hit) return null;
if ((Date.now() - hit.ts) >= SUGGEST_CACHE_TTL_MS) {
suggestCache.delete(key);
return null;
}
// LRU refresh
suggestCache.delete(key);
suggestCache.set(key, hit);
return hit.data;
}
function suggestCacheSet(key, data) {
if (suggestCache.has(key)) suggestCache.delete(key);
suggestCache.set(key, { ts: Date.now(), data });
while (suggestCache.size > SUGGEST_CACHE_MAX_ENTRIES) {
const oldest = suggestCache.keys().next().value;
suggestCache.delete(oldest);
}
}
const suggestLimiter = rateLimit({
windowMs: 60 * 1000,
max: Number(process.env.SUGGEST_RATE_LIMIT || 60),
standardHeaders: true,
legacyHeaders: false,
});
app.get('/api/search/suggest', suggestLimiter, async (req, res) => {
try {
const rawQ = typeof req.query.q === 'string' ? req.query.q : '';
const q = rawQ.trim();
if (q.length < 2) {
return res.status(400).json({ error: 'q is required and must be at least 2 characters long' });
}
const limit = Math.min(20, Math.max(1, Number(req.query.limit || 10)));
const requested = typeof req.query.providers === 'string' ? String(req.query.providers) : '';
const validProviders = validateProviders(requested);
const cacheKey = `suggest:${validProviders.join(',')}:${q.toLowerCase()}:${limit}`;
const cached = suggestCacheGet(cacheKey);
if (cached) return res.json(cached);
const results = await Promise.allSettled(
validProviders.map((providerId) => {
const mod = providerRegistry[providerId];
if (!mod || typeof mod.suggest !== 'function') return Promise.resolve([]);
return Promise.resolve().then(() => mod.suggest(q, { limit }));
})
);
const groups = {};
results.forEach((result, index) => {
const providerId = validProviders[index];
if (result.status === 'fulfilled' && Array.isArray(result.value)) {
groups[providerId] = result.value
.map((s) => String(s ?? '').trim())
.filter(Boolean)
.slice(0, limit);
} else {
groups[providerId] = [];
}
});
const data = { q, groups };
suggestCacheSet(cacheKey, data);
return res.json(data);
} catch (e) {
return res.status(500).json({ error: 'suggest_failed', details: String(e?.message || e) });
}
});
// -------------------- Video transcripts (Step 16, Phase 1) --------------------
// GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=
// -> { lang, available, languages, lines: [{ t, dur, text }] }
// One endpoint / one parser / one UI whatever the provider (yt-dlp subtitles).
const TRANSCRIPT_CACHE_TTL_MS = Number(process.env.TRANSCRIPT_CACHE_TTL || 24 * 60 * 60 * 1000);
const TRANSCRIPT_CACHE_MAX_ENTRIES = 200;
/** @type {Map<string, { ts: number, data: any }>} */
const transcriptCache = new Map();
function transcriptCacheGet(key) {
const hit = transcriptCache.get(key);
if (!hit) return null;
if ((Date.now() - hit.ts) >= TRANSCRIPT_CACHE_TTL_MS) {
transcriptCache.delete(key);
return null;
}
transcriptCache.delete(key);
transcriptCache.set(key, hit);
return hit.data;
}
function transcriptCacheSet(key, data) {
if (transcriptCache.has(key)) transcriptCache.delete(key);
transcriptCache.set(key, { ts: Date.now(), data });
while (transcriptCache.size > TRANSCRIPT_CACHE_MAX_ENTRIES) {
const oldest = transcriptCache.keys().next().value;
transcriptCache.delete(oldest);
}
}
const transcriptLimiter = rateLimit({
windowMs: 60 * 1000,
max: Number(process.env.TRANSCRIPT_RATE_LIMIT || 10),
standardHeaders: true,
legacyHeaders: false,
// Always answer JSON (default handler sends an HTML/text page, which the
// Angular HttpClient cannot parse as JSON and surfaces as a raw SyntaxError).
handler: (req, res) => {
return res.status(429).json({ available: false, error: 'rate_limited' });
},
});
/** Providers known NOT to expose subtitle tracks via yt-dlp (no transcript possible). */
const TRANSCRIPT_UNSUPPORTED_PROVIDERS = new Set(['twitch', 'odysee', 'rumble']);
/** Download subtitles via yt-dlp (handles YouTube impersonation + 429-prone
* translated tracks). Tries `langs` in order, returns the first non-empty
* parsed lines with the language that worked, or null. Bounded by a timeout
* (yt-dlp subtitle downloads can hang on .part files); partial results on
* disk are still scanned when yt-dlp exits non-zero. */
async function transcriptViaYtDlp(url, langs) {
const osMod = await import('node:os');
const dir = await fs.promises.mkdtemp(path.join(osMod.tmpdir(), 'newtube-transcript-'));
const scanDir = async (wanted) => {
let files = [];
try {
files = (await fs.promises.readdir(dir)).filter((f) => /\.vtt$/i.test(f) && !/\.part$/i.test(f));
} catch { return null; }
// Ignore stale/empty files
const nonEmpty = [];
for (const f of files) {
try {
const st = await fs.promises.stat(path.join(dir, f));
if (st.size > 0) nonEmpty.push(f);
} catch {}
}
// Prefer requested languages first
nonEmpty.sort((a, b) => {
const la = a.toLowerCase(), lb = b.toLowerCase();
const ia = wanted.findIndex((w) => la.includes(`.${w.toLowerCase()}.`) || la.endsWith(`.${w.toLowerCase()}.vtt`));
const ib = wanted.findIndex((w) => lb.includes(`.${w.toLowerCase()}.`) || lb.endsWith(`.${w.toLowerCase()}.vtt`));
return (ia === -1 ? 99 : ia) - (ib === -1 ? 99 : ib);
});
for (const file of nonEmpty) {
try {
const text = await fs.promises.readFile(path.join(dir, file), 'utf8');
const lines = parseVtt(text);
if (lines.length) {
const m = /\.([a-z]{2,3}(?:-[a-z]{2,4})?)\.vtt$/i.exec(file);
return { lines, lang: m ? m[1] : null };
}
} catch {}
}
return null;
};
try {
const wanted = Array.from(new Set((langs || []).map((l) => String(l || '').split('-')[0]).filter(Boolean)));
if (!wanted.includes('en')) wanted.push('en');
const subLangs = wanted.slice(0, 4).join(',');
const child = youtubedl(url, {
writeSub: true,
writeAutoSub: true,
subLangs,
subFormat: 'vtt/best',
skipDownload: true,
noWarnings: true,
noCheckCertificates: true,
noPlaylist: true,
output: path.join(dir, '%(id)s'),
});
const timer = setTimeout(() => { try { child.kill('SIGKILL'); } catch {} }, 90000);
try {
await child;
} catch (e) {
// Non-zero exit (e.g. one language 429'd) — partial files may still exist.
console.warn('[transcript] yt-dlp subtitle download exited non-zero:', String(e?.message || e).slice(0, 200));
} finally {
clearTimeout(timer);
}
return await scanDir(wanted);
} catch (e) {
console.warn('[transcript] yt-dlp subtitle download failed:', e?.message || e);
try {
const wanted = Array.from(new Set((langs || []).map((l) => String(l || '').split('-')[0]).filter(Boolean)));
return await scanDir(wanted);
} catch { return null; }
} finally {
try { await fs.promises.rm(dir, { recursive: true, force: true }); } catch {}
}
}
/** Fetch one timedtext track URL and parse it (any format). Returns [] on any failure. */
async function fetchTimedTextLines(track) {
const resp = await fetch(String(track.url), {
headers: {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36',
Accept: 'application/json, text/vtt, text/*;q=0.9, */*;q=0.8',
'Accept-Language': 'en-US,en;q=0.9,fr;q=0.8',
},
signal: AbortSignal.timeout(15000),
});
if (!resp.ok) throw new Error(`track_fetch_failed:${resp.status}`);
const text = await resp.text();
if (!text || looksLikeHtmlError(text)) throw new Error('track_fetch_failed:html_error_page');
return parseTrackText(text, transcriptTrackExt(track));
}
app.get('/api/transcript/:provider/:videoId', transcriptLimiter, async (req, res) => {
try {
const { provider, videoId } = req.params;
const lang = String(req.query.lang || 'fr').slice(0, 12) || 'fr';
const normalized = normalizeTranscriptProvider(provider);
if (!normalized || !videoId) {
return res.status(400).json({ available: false, error: 'invalid_provider_or_video' });
}
const cacheKey = `transcript:${normalized}:${videoId}:${lang.toLowerCase()}`;
const cached = transcriptCacheGet(cacheKey);
if (cached) return res.json(cached);
// Source de découverte des pistes : InnerTube d'abord pour YouTube
// (mêmes URLs timedtext, sans spawn yt-dlp), yt-dlp sinon/en secours.
// YT_TRANSCRIPT_SOURCE=innertube-first (défaut) | ytdlp-only | innertube-only
const transcriptSource = String(process.env.YT_TRANSCRIPT_SOURCE || 'innertube-first').trim().toLowerCase();
const transcriptUrl = providerUrlFrom(normalized, String(videoId), {
instance: req.query.instance || undefined,
slug: req.query.slug || undefined,
sourceUrl: req.query.sourceUrl || undefined,
});
let meta = null;
if (normalized === 'youtube' && transcriptSource !== 'ytdlp-only') {
try {
const { getCaptionTracksViaInnerTube } = await import('./providers/youtube-innertube.mjs');
const cap = await getCaptionTracksViaInnerTube(String(videoId));
if (cap.trackCount > 0) {
meta = { subtitles: cap.subtitles, automatic_captions: cap.automatic_captions };
console.log(`[transcript] source=innertube tracks=${cap.trackCount} langs=${cap.languages.join(',')}`);
} else {
// InnerTube fait foi (même backend que le lecteur) : 0 piste = pas
// de sous-titres, sans payer un dump yt-dlp complet.
console.log('[transcript] source=innertube tracks=0 -> no_subtitles');
const empty = { lang: null, available: false, languages: [], lines: [], reason: 'no_subtitles' };
transcriptCacheSet(cacheKey, empty);
return res.json(empty);
}
} catch (e) {
console.warn('[transcript] innertube discovery failed, fallback yt-dlp:', e?.code || String(e?.message || e).slice(0, 120));
if (transcriptSource === 'innertube-only') {
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
}
}
}
if (!meta) {
try {
const raw = await youtubedl(transcriptUrl, { dumpSingleJson: true, skipDownload: true, noWarnings: true, noCheckCertificates: true });
meta = (typeof raw === 'string') ? JSON.parse(raw || '{}') : (raw || {});
} catch (e) {
console.error('[transcript] yt-dlp failed:', e?.message || e);
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
}
}
const { track, languages, lang: chosenLang } = pickTrack(meta, lang);
if (!track) {
// Definitive absence: distinguish "provider never exposes subtitles"
// (twitch/odysee/rumble) from "this video has none" — cacheable 200s.
const reason = TRANSCRIPT_UNSUPPORTED_PROVIDERS.has(normalized) ? 'provider_unsupported' : 'no_subtitles';
const empty = { lang: null, available: false, languages: languages || [], lines: [], reason };
transcriptCacheSet(cacheKey, empty);
return res.json(empty);
}
// Try candidate tracks in order: requested language, then original (`en`),
// then everything else. Translated tracks are often rate-limited while the
// original still works — never fail on the first track alone.
// Bounded: hammering dozens of timedtext URLs only worsens YouTube 429s.
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
let lines = [];
let workedLang = chosenLang;
let sawRateLimit = false;
for (const cand of orderedTracks(meta, lang).slice(0, 5)) {
let attempt = 0;
for (;;) {
try {
const parsed = await fetchTimedTextLines(cand.track);
if (parsed && parsed.length) {
lines = parsed;
workedLang = cand.lang || chosenLang;
}
break;
} catch (e) {
const msg = String(e?.message || e);
console.warn('[transcript] track fetch failed:', cand.lang, msg);
// One retry after a short pause on transient 429s.
if (msg.includes(':429') && attempt === 0) {
sawRateLimit = true;
attempt += 1;
await sleep(1500);
continue;
}
if (msg.includes(':429')) sawRateLimit = true;
break;
}
}
if (lines.length) break;
}
if (!lines || lines.length === 0) {
// Direct timedtext fetches failed (429 / Sorry pages / impersonation):
// let yt-dlp download the subtitles instead (handles all of the above).
try {
const viaDlp = await transcriptViaYtDlp(transcriptUrl, [lang, chosenLang].filter(Boolean));
if (viaDlp && viaDlp.lines && viaDlp.lines.length) {
lines = viaDlp.lines;
if (viaDlp.lang) workedLang = viaDlp.lang;
}
} catch {}
}
if (!lines || lines.length === 0) {
// Subtitle tracks exist but their content could not be retrieved or
// parsed (YouTube 429 / Sorry pages / impersonation): the video HAS
// subtitles, so this is transient — 502 with languages + retryable flag,
// never cached as "no subtitles".
return res.status(502).json({
available: false,
languages: languages || [],
lines: [],
error: 'transcript_temporarily_unavailable',
retryable: true,
rateLimited: sawRateLimit || undefined,
});
}
const result = { lang: workedLang, available: true, languages: languages || [], lines };
transcriptCacheSet(cacheKey, result);
return res.json(result);
} catch (e) {
console.error('[transcript] unexpected error:', e?.message || e);
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
}
});
// -------------------- Static Frontend (Angular build) --------------------
const distRoot = path.join(process.cwd(), 'dist');
const distBrowser = path.join(distRoot, 'browser');
+34 -1
View File
@@ -63,6 +63,16 @@ async function resolveYouTubeChannelId(externalId) {
const raw = String(externalId || '').trim();
if (/^UC[\w-]{20,}$/.test(raw)) return raw;
const handle = raw.replace(/^@/, '');
// 0) scrape sans clé (Step 17) : ne consomme aucun quota
try {
const { getSearchMode } = await import('./youtube-common.mjs');
const mode = getSearchMode();
if (mode !== 'api-only') {
const { resolveChannelIdViaScrape } = await import('./youtube-scrape.mjs');
const id = await resolveChannelIdViaScrape(raw);
if (id && /^UC[\w-]{20,}$/.test(id)) return id;
}
} catch {}
// 1) channels?forHandle (fonctionne encore pour beaucoup de chaînes)
try {
const data = await ytGet('channels', { part: 'id', forHandle: handle });
@@ -131,9 +141,32 @@ function ytTokenFor(key, page) {
}
async function ytContent(externalId, { type, page, limit, sort, q }) {
const channelId = await resolveYouTubeChannelId(externalId);
const perPage = Math.min(Math.max(1, Number(limit || 24)), 50);
const pageNum = Math.max(1, Number(page || 1));
// Step 17 : scrape-first sans clé (0 quota). Fallback API si bot-check/timeout.
try {
const { getSearchMode } = await import('./youtube-common.mjs');
const mode = getSearchMode();
if (mode !== 'api-only') {
const { fetchChannelViaScrape } = await import('./youtube-scrape.mjs');
// externalId brut (handle ou UC...) : le scrape gère les deux formes
const scraped = await fetchChannelViaScrape(externalId, { type, page: pageNum, limit: perPage });
if (Array.isArray(scraped?.items) && scraped.items.length) {
if (mode === 'scrape-only') return { ...scraped, total: null };
// scrape-first : retour direct si non vide
return { ...scraped, total: null };
}
// vide -> on tente l'API (chaîne à faible volume ou tab non supporté en scrape)
if (mode === 'scrape-only') return { items: [], nextPage: null };
}
} catch (e) {
console.warn('[channel-content/yt] scrape failed, fallback api:', e?.code || e?.message || e);
try {
const { getSearchMode } = await import('./youtube-common.mjs');
if (getSearchMode() === 'scrape-only') return { items: [], nextPage: null };
} catch {}
}
const channelId = await resolveYouTubeChannelId(externalId);
if (type === 'playlists') {
const key = ['pl', channelId, perPage].join('|');
const token = ytTokenFor(key, pageNum);
+174
View File
@@ -0,0 +1,174 @@
// Helpers partagés YouTube : clés API, mode de recherche, anti-ban yt-dlp, métriques.
// Extrait de youtube.mjs + index.mjs pour éviter la duplication (Step 17 P0).
import fs from 'node:fs';
import path from 'node:path';
import crypto from 'node:crypto';
import { execFile as execFileCb } from 'node:child_process';
import { promisify } from 'node:util';
export function getYouTubeKeys() {
const keys = [];
try {
const raw = process.env.YOUTUBE_API_KEYS;
if (raw && String(raw).trim() && !['undefined', 'null'].includes(String(raw).trim())) {
const s = String(raw).trim();
if (s.startsWith('[')) {
try {
const arr = JSON.parse(s);
if (Array.isArray(arr)) keys.push(...arr.map((v) => String(v || '').trim()).filter(Boolean));
} catch {}
} else {
keys.push(...s.split(',').map((v) => String(v || '').trim()).filter(Boolean));
}
}
} catch {}
try {
const single = process.env.YOUTUBE_API_KEY;
if (single && String(single).trim()) keys.push(String(single).trim());
} catch {}
return Array.from(new Set(keys.filter(Boolean)));
}
export function isKeyFailure(status, data) {
try {
const reason = data?.error?.errors?.[0]?.reason || '';
const message = String(data?.error?.message || '');
if (status === 400 && (reason === 'API_KEY_INVALID' || /api key (expired|invalid)/i.test(message))) return true;
if (status === 403 && /quota|rateLimit|dailyLimit|userRateLimit/i.test(`${reason} ${message}`)) return true;
} catch {}
return false;
}
/** Mode de recherche YT : innertube-first (défaut) | innertube-only | scrape-first | api-first | scrape-only | api-only */
export const YT_SEARCH_MODES = ['innertube-first', 'innertube-only', 'scrape-first', 'api-first', 'scrape-only', 'api-only'];
export function getSearchMode() {
const m = String(process.env.YT_SEARCH_MODE || 'innertube-first').trim().toLowerCase();
if (YT_SEARCH_MODES.includes(m)) return m;
return 'innertube-first';
}
export function getScrapeTtlMs() {
const v = Number(process.env.YT_SCRAPE_TTL_MS || 30 * 60 * 1000);
return Number.isFinite(v) && v > 0 ? v : 30 * 60 * 1000;
}
export function getYtDlpTimeoutMs() {
const v = Number(process.env.YT_DLP_TIMEOUT_MS || 20000);
return Number.isFinite(v) && v >= 5000 ? v : 20000;
}
export function getYtDlpBin() {
const p = String(process.env.YT_DLP_PATH || '').trim();
if (p) return p;
return process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
}
/**
* Résolution robuste du binaire yt-dlp (même ordre que server/index.mjs) :
* 1. YT_DLP_PATH (si le fichier existe)
* 2. `yt-dlp[.exe]` sur le PATH (probé via --version)
* 3. binaire bundled de `youtube-dl-exec` (node_modules/.../bin/, utilisé en Docker)
* Résultat mis en cache process-wide. Throw avec code `yt_scrape_no_binary` si introuvable.
*/
let _resolvedBin = null;
export function resetYtDlpBinCache() { _resolvedBin = null; }
function bundledYtDlpPath() {
try {
const exe = process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
// server/providers/ -> server/../node_modules/youtube-dl-exec/bin/
const cand = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, '$1')), '..', '..', 'node_modules', 'youtube-dl-exec', 'bin', exe);
if (fs.existsSync(cand)) return cand;
} catch {}
return null;
}
export function getBundledYtDlpPath() { return bundledYtDlpPath(); }
async function probeBin(bin) {
try {
const execFileAsync = promisify(execFileCb);
await execFileAsync(bin, ['--version'], { timeout: 10000 });
return true;
} catch (e) {
// ENOENT = binaire absent ; autre erreur (timeout...) = présent mais KO -> on le garde quand même
return e?.code !== 'ENOENT' && e?.errno !== 'ENOENT';
}
}
export async function resolveYtDlpBin() {
if (_resolvedBin) return _resolvedBin;
// 1) YT_DLP_PATH explicite
try {
const p = String(process.env.YT_DLP_PATH || '').trim();
if (p && fs.existsSync(p)) {
if (await probeBin(p)) { _resolvedBin = p; return p; }
} else if (p) {
console.warn(`[YT] YT_DLP_PATH introuvable : ${p} (fallback PATH/bundled)`);
}
} catch {}
// 2) PATH
const exe = process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
for (const cand of [exe, 'yt-dlp']) {
try {
if (await probeBin(cand)) { _resolvedBin = cand; return cand; }
} catch {}
}
// 3) bundled youtube-dl-exec
const bundled = bundledYtDlpPath();
if (bundled && await probeBin(bundled)) { _resolvedBin = bundled; return bundled; }
throw Object.assign(
new Error('yt-dlp introuvable (PATH, YT_DLP_PATH ni binaire bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
);
}
/**
* Args anti-ban communs pour yt-dlp (cookies, PO-Token, proxy).
* Ne jamais logger les valeurs (cookies path exclu des logs verbeux).
*/
export function buildYtDlpExtraArgs() {
const args = [];
try {
const cookies = String(process.env.YT_COOKIES_FILE || '').trim();
if (cookies && fs.existsSync(cookies)) {
args.push('--cookies', cookies);
}
} catch {}
try {
const po = String(process.env.YT_PO_TOKEN || '').trim();
if (po) args.push('--extractor-args', `youtube:po_token=${po}`);
} catch {}
try {
const proxy = String(process.env.YT_EGRESS_PROXY || '').trim();
if (proxy) args.push('--proxy', proxy);
} catch {}
return args;
}
export function hasCookiesFile() {
try {
const c = String(process.env.YT_COOKIES_FILE || '').trim();
return Boolean(c && fs.existsSync(c));
} catch { return false; }
}
export function hashSearchKey(parts) {
return crypto.createHash('sha256').update(String(parts)).digest('hex').slice(0, 32);
}
// Métriques process-wide (remises à zéro au restart, persistées jour par jour en SQLite via db.mjs)
export const ytMetrics = {
scrapeCalls: 0,
scrapeHits: 0, // hit cache (mémoire ou SQLite)
apiCalls: 0,
quotaUnits: 0, // estimation search*100 + videos*1
fallbacks: 0,
scrapeErrors: 0,
innertubeCalls: 0,
innertubeErrors: 0,
};
export function metricsSnapshot() {
return { ...ytMetrics };
}
+331
View File
@@ -0,0 +1,331 @@
// Couche InnerTube directe (façon SmartTube/MediaServiceCore) via youtubei.js.
// WEB client : search + continuations (pagination illimitée), watch-next (related),
// sans clé API ni quota. Le dispatcher bascule sur scrape/API si indisponible.
import { hashSearchKey, ytMetrics } from './youtube-common.mjs';
let sessionPromise = null;
let LogSilenced = false;
async function silenceLibNoise() {
if (LogSilenced) return;
LogSilenced = true;
try {
const { Log } = await import('youtubei.js');
// youtubei.js loggue en WARN chaque Text sans run assorti (bruit sur les titres) :
// on ne garde que les erreurs.
if (Log?.set_level && Log?.Level) {
const lvl = Log.Level.ERROR ?? Log.Level.WARNING ?? 3;
Log.set_level(lvl);
}
} catch {}
}
/** Session InnerTube singleton (lazy). Throw code yt_innertube_unavailable si KO. */
export async function getSession() {
if (!sessionPromise) {
sessionPromise = (async () => {
try {
await silenceLibNoise();
const { Innertube } = await import('youtubei.js');
const gl = String(process.env.YT_INNERTUBE_GL || 'FR').trim() || 'FR';
const hl = String(process.env.YT_INNERTUBE_HL || 'fr').trim() || 'fr';
return await Innertube.create({ lang: hl, location: gl });
} catch (e) {
sessionPromise = null; // retry au prochain appel
throw Object.assign(
new Error(`InnerTube indisponible : ${String(e?.message || e).slice(0, 160)}`),
{ ytStatus: 502, code: 'yt_innertube_unavailable' },
);
}
})();
}
return sessionPromise;
}
export function resetSession() { sessionPromise = null; }
/** "81 973 vues" / "57K" / "1,2 M vues" -> number | undefined. */
export function parseViewsText(t) {
try {
const s = String(t || '').trim();
if (!s) return undefined;
const m = s.match(/([\d\s\u00a0.,]+)\s*([KMBkmb]|Mds|M|k|B)?/);
if (!m) return undefined;
const num = Number(m[1].replace(/[\s\u00a0]/g, '').replace(',', '.'));
if (!Number.isFinite(num)) return undefined;
const suffix = (m[2] || '').toLowerCase();
const mult = suffix === 'b' ? 1e9 : suffix === 'm' || suffix === 'mds' ? 1e6 : suffix === 'k' ? 1e3 : 1;
const v = Math.round(num * mult);
return v > 0 ? v : undefined;
} catch { return undefined; }
}
/** "2 minutes, 27 seconds" / "1 heure, 5 minutes" (label a11y) -> secondes. */
export function parseDurationLabel(label) {
try {
const s = String(label || '').toLowerCase();
if (!s) return undefined;
const get = (re) => { const m = s.match(re); return m ? Number(m[1]) : 0; };
const h = get(/(\d+)\s*(?:hours?|heures?)/);
const mnt = get(/(\d+)\s*(?:minutes?)/);
const sec = get(/(\d+)\s*(?:seconds?|secondes?)/);
const total = h * 3600 + mnt * 60 + sec;
return total > 0 ? total : undefined;
} catch { return undefined; }
}
function bestThumb(thumbs) {
try {
const arr = Array.isArray(thumbs) ? thumbs.filter((t) => t?.url) : [];
if (!arr.length) return undefined;
return arr[arr.length - 1].url;
} catch { return undefined; }
}
/**
* Mappe un node InnerTube (Video, CompactVideo, GridVideo, LockupView, ReelItem,
* ShortsLockupView, PlaylistPanelVideo, WatchCardCompactVideo) vers Suggestion.
* Pur et testable offline. Les LockupView (nouveau renderer YouTube, utilisé
* notamment dans le watch-next façon SmartTube) ont une forme imbriquée propre.
*/
export function mapVideoNode(n) {
if (!n || typeof n !== 'object') return null;
try {
const type = String(n.type || '');
if (type === 'LockupView') return mapLockupView(n);
if (/playlist|channel|gridchannel|shelf|radio|show|album/i.test(type)
&& !/playlistpanelvideo|watchcard/i.test(type)) return null;
const id = n.id ? String(n.id) : null;
if (!id) return null;
const title = n.title?.text ?? (typeof n.title === 'string' ? n.title : '');
if (!title) return null;
const duration = n.duration?.seconds != null ? Number(n.duration.seconds) : undefined;
const author = n.author || n.uploader || null;
const authorName = author?.name ?? (typeof author === 'string' ? author : undefined);
const authorId = author?.id ? String(author.id) : undefined;
const views = parseViewsText(n.view_count?.text ?? n.view_count ?? n.views?.text);
const isShort = /short|reel/i.test(type) || (Number.isFinite(duration) && duration > 0 && duration <= 70 && /short/i.test(title) === false && /reel|short/i.test(type));
const badges = Array.isArray(n.badges) ? n.badges.map((b) => b?.label).filter(Boolean) : [];
return {
title: String(title),
id,
url: `https://www.youtube.com/watch?v=${id}`,
thumbnail: bestThumb(n.thumbnails),
uploaderName: authorName,
type: 'video',
...(Number.isFinite(duration) && duration > 0 ? { duration } : {}),
...(views !== undefined ? { views } : {}),
...(n.published?.text ? { publishedAt: String(n.published.text) } : {}),
...(authorId ? { channelId: authorId, channelExternalId: authorId, channelUrl: `https://www.youtube.com/channel/${authorId}` } : {}),
...(authorName ? { channelHandle: String(authorName) } : {}),
...(n.is_live ? { isLive: true } : {}),
...(isShort ? { isShort: true } : {}),
...(badges.length ? { badges } : {}),
};
} catch { return null; }
}
/** Mappe un LockupView (renderer moderne : watch-next, search, shelves). */
export function mapLockupView(n) {
try {
const ct = String(n.content_type || '').toUpperCase();
if (ct && !['VIDEO', 'SHORT', 'MOVIE', 'LIVE'].includes(ct)) return null;
const id = n.content_id ? String(n.content_id)
: (n.renderer_context?.command_context?.on_tap?.payload?.videoId || null);
if (!id) return null;
const meta = n.metadata || {};
const title = meta.title?.text ?? (typeof meta.title === 'string' ? meta.title : '');
if (!title) return null;
// Vignette : content_image.image[] (prend la plus large)
let thumbnail;
try {
const imgs = (n.content_image?.image || []).filter((i) => i?.url);
thumbnail = imgs.sort((a, b) => (b.width || 0) - (a.width || 0))[0]?.url;
} catch {}
// Chaîne : 1ère ligne des metadata_rows, sinon label a11y "Go to channel X"
let uploaderName;
try {
const rows = meta.metadata?.metadata_rows || [];
const first = rows[0]?.metadata_parts?.[0]?.text;
uploaderName = first?.text ?? (typeof first === 'string' ? first : undefined);
} catch {}
if (!uploaderName) {
const a11y = String(meta.image?.a11y_label || '');
const m = a11y.match(/^(?:go to channel|aller sur la cha[îi]ne)\s+(.+)$/i);
if (m) uploaderName = m[1].trim();
}
// Vues : 2e ligne ("57K", "1,2 M vues"...), durée : label a11y
let views;
try {
const rows = meta.metadata?.metadata_rows || [];
const second = rows[1]?.metadata_parts?.[0]?.text;
views = parseViewsText(second?.text ?? second);
} catch {}
const duration = parseDurationLabel(n.renderer_context?.accessibility_context?.label);
return {
title: String(title),
id,
url: `https://www.youtube.com/watch?v=${id}`,
thumbnail,
uploaderName,
type: 'video',
...(duration !== undefined ? { duration } : {}),
...(views !== undefined ? { views } : {}),
...(uploaderName ? { channelHandle: String(uploaderName) } : {}),
...(ct === 'SHORT' ? { isShort: true } : {}),
...(ct === 'LIVE' ? { isLive: true } : {}),
};
} catch { return null; }
}
export function mapNodes(nodes) {
const seen = new Set();
const out = [];
for (const n of nodes || []) {
const m = mapVideoNode(n);
if (!m || seen.has(m.id)) continue;
seen.add(m.id);
out.push(m);
}
return out;
}
// Chaîne de continuations en mémoire : q_hash -> { feed, pages }.
// (SQLite persiste les pages déjà servies ; la chaîne évite de rejouer les pages 1..N-1.)
const chainCache = new Map();
const CHAIN_MAX = 50;
function chainGet(k) { const h = chainCache.get(k); if (h) { chainCache.delete(k); chainCache.set(k, h); } return h || null; }
function chainSet(k, v) {
if (chainCache.has(k)) chainCache.delete(k);
chainCache.set(k, v);
while (chainCache.size > CHAIN_MAX) { const o = chainCache.keys().next().value; if (o === undefined) break; chainCache.delete(o); }
}
export function innertubeChainStats() { return { chains: chainCache.size, max: CHAIN_MAX }; }
function sortParam(sort) {
const s = String(sort || 'relevance').toLowerCase();
if (s === 'date') return 'upload_date';
if (s === 'views') return 'view_count';
return 'relevance';
}
/**
* Recherche InnerTube avec vraie pagination (continuations).
* Page 1 ~35-40 vidéos ; pages suivantes via getContinuation() fusionnées (façon SmartTube).
*/
export async function searchViaInnerTube(q, opts = {}) {
const query = String(q || '').trim();
if (query.length < 2) return [];
const limit = Math.min(50, Math.max(1, Number(opts?.limit || 24)));
const page = Math.min(10, Math.max(1, Number(opts?.page || 1)));
const sort = String(opts?.sort || 'relevance');
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
const yt = await getSession();
const key = `it|${hashSearchKey(`${query.toLowerCase()}|${sort}`)}`;
let entry = chainGet(key);
let feed = entry?.feed || null;
let collected = entry?.items ? [...entry.items] : [];
let pagesDone = entry?.pages || 0;
// Note : une "page" InnerTube fait ~17-20 vidéos quel que soit `limit`.
// On charge donc des continuations jusqu'à couvrir la fenêtre demandée
// [start, start+limit[ (façon SmartTube qui remplit son écran au fil des
// continuations), au lieu d'aligner 1 page API = 1 page UI.
const need = page * limit;
try {
if (!feed) {
feed = await yt.search(query, { sort_by: sortParam(sort) });
collected = mapNodes(feed.videos);
pagesDone = 1;
}
let guard = 0;
while (collected.length < need && feed?.has_continuation && guard < 12) {
guard++;
feed = await feed.getContinuation();
const fresh = mapNodes(feed.videos).filter((m) => !collected.some((c) => c.id === m.id));
collected = collected.concat(fresh);
pagesDone++;
if (!fresh.length) break;
}
} catch (e) {
if (e?.code === 'yt_innertube_unavailable') throw e;
throw Object.assign(new Error(`InnerTube search failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
}
if (feed) chainSet(key, { feed, items: collected, pages: pagesDone });
const start = (page - 1) * limit;
return collected.slice(start, start + limit);
}
/**
* Vidéos connexes façon SmartTube (endpoint watch-next) : ce qui alimente
* la colonne "À suivre" sous le lecteur.
*/
export async function getRelatedViaInnerTube(videoId, limit = 24) {
const id = String(videoId || '').trim();
if (!id) return [];
const n = Math.min(50, Math.max(1, Number(limit || 24)));
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
const yt = await getSession();
try {
const info = await yt.getInfo(id);
const feed = info?.watch_next_feed;
const arr = Array.isArray(feed) ? feed : (feed ? Array.from(feed) : []);
return mapNodes(arr).filter((m) => m.id !== id).slice(0, n);
} catch (e) {
if (e?.code === 'yt_innertube_unavailable') throw e;
throw Object.assign(new Error(`InnerTube related failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
}
}
// -------------------- Transcripts via InnerTube --------------------
// `getInfo().captions.caption_tracks[]` expose les mêmes URLs timedtext que
// yt-dlp découvre via un dump complet (`subtitles`/`automatic_captions`).
// On adapte leur forme vers le format yt-dlp pour réutiliser pickTrack(),
// orderedTracks() et parseTrackText() de transcript.mjs sans les toucher.
function normCaptionLang(code) {
return String(code || '').trim().toLowerCase().replace(/_/g, '-');
}
/**
* Adapte des caption tracks InnerTube brutes vers un pseudo dump yt-dlp
* `{ subtitles, automatic_captions }` (kind 'asr' = auto-généré).
* Pur et testable offline.
*/
export function mapCaptionTracks(captionTracks) {
const subtitles = {};
const automatic_captions = {};
for (const t of captionTracks || []) {
const base = String(t?.base_url || '');
if (!base) continue;
const lang = normCaptionLang(t?.language_code) || 'und';
const name = t?.name?.text ?? (typeof t?.name === 'string' ? t.name : lang);
const sep = base.includes('?') ? '&' : '?';
const track = { url: `${base}${sep}fmt=json3`, ext: 'json3', name: String(name) };
const dict = t?.kind === 'asr' ? automatic_captions : subtitles;
if (!Array.isArray(dict[lang])) dict[lang] = [];
if (!dict[lang].some((x) => x.url === track.url)) dict[lang].push(track);
}
const languages = Array.from(new Set([...Object.keys(subtitles), ...Object.keys(automatic_captions)]));
const trackCount = Object.values(subtitles).concat(Object.values(automatic_captions))
.reduce((n, arr) => n + arr.length, 0);
return { subtitles, automatic_captions, languages, trackCount };
}
/**
* Découvre les pistes de sous-titres d'une vidéo via InnerTube (0 quota,
* pas de spawn yt-dlp). Retourne un pseudo dump yt-dlp prêt pour pickTrack().
*/
export async function getCaptionTracksViaInnerTube(videoId) {
const id = String(videoId || '').trim();
if (!id) return { subtitles: {}, automatic_captions: {}, languages: [], trackCount: 0 };
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
const yt = await getSession();
try {
const info = await yt.getInfo(id);
const raw = info?.captions?.caption_tracks || [];
return mapCaptionTracks(Array.isArray(raw) ? raw : Array.from(raw));
} catch (e) {
if (e?.code === 'yt_innertube_unavailable') throw e;
throw Object.assign(new Error(`InnerTube captions failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
}
}
+197
View File
@@ -0,0 +1,197 @@
// Recherche / channel / trending YouTube SANS clé API via yt-dlp (Step 17 P1).
// Binaire résolu via YT_DLP_PATH > PATH. Jamais de secret en log.
import { execFile as execFileCb } from 'node:child_process';
import { promisify } from 'node:util';
import { getYtDlpTimeoutMs, buildYtDlpExtraArgs, resolveYtDlpBin } from './youtube-common.mjs';
const execFileAsync = promisify(execFileCb);
/**
* Normalise une entrée yt-dlp flat-playlist vers Suggestion (même contrat que youtube.mjs).
* @param {any} e entrée yt-dlp
*/
export function mapFlatEntry(e) {
if (!e || typeof e !== 'object') return null;
const id = e.id || e.display_id || null;
if (!id) return null;
const thumbs = Array.isArray(e.thumbnails) ? e.thumbnails : [];
const thumb = thumbs.length
? (thumbs.find((t) => t?.url && (t.height || 0) >= 360)?.url || thumbs[thumbs.length - 1]?.url)
: (e.thumbnail || undefined);
const duration = e.duration != null ? Number(e.duration) : undefined;
const views = e.view_count != null ? Number(e.view_count) : (e.views != null ? Number(e.views) : undefined);
const channelId = e.channel_id || e.uploader_id || undefined;
const uploaderName = e.channel || e.uploader || undefined;
return {
title: e.title || '',
id: String(id),
url: String(id).startsWith('http') ? String(id) : `https://www.youtube.com/watch?v=${id}`,
thumbnail: thumb,
uploaderName,
type: 'video',
...(Number.isFinite(duration) && duration > 0 ? { duration } : {}),
...(Number.isFinite(views) ? { views } : {}),
...(e.timestamp ? { publishedAt: new Date(Number(e.timestamp) * 1000).toISOString() } : {}),
...(channelId ? { channelId: String(channelId), channelExternalId: String(channelId), channelUrl: `https://www.youtube.com/channel/${channelId}` } : {}),
...(uploaderName ? { channelHandle: String(uploaderName) } : {}),
};
}
/** Parse le JSON de `yt-dlp --dump-single-json --flat-playlist` (objet ou NDJSON fog). */
export function parseFlatPlaylistJson(raw) {
try {
const text = String(raw || '').trim();
if (!text) return [];
// Cas standard : un seul objet JSON avec .entries
try {
const obj = JSON.parse(text);
const entries = Array.isArray(obj?.entries) ? obj.entries : (Array.isArray(obj) ? obj : []);
return entries.map(mapFlatEntry).filter(Boolean);
} catch {
// Fallback NDJSON (une ligne = un JSON)
const out = [];
for (const line of text.split('\n')) {
const t = line.trim();
if (!t) continue;
try {
const o = JSON.parse(t);
const m = mapFlatEntry(o);
if (m) out.push(m);
} catch {}
}
return out;
}
} catch { return []; }
}
export function classifyScrapeError(e) {
if (e?.code === 'yt_scrape_no_binary') return e;
const msg = String(e?.message || e || '');
if (/ENOENT/i.test(msg) || e?.code === 'ENOENT' || e?.errno === 'ENOENT') {
return Object.assign(
new Error('yt-dlp introuvable (ni PATH ni bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
);
}
if (/bot|sign in to confirm|challenge|cookies|login/i.test(msg)) {
return Object.assign(new Error('YouTube bot-check (cookies/PO-Token requis)'), { ytStatus: 502, code: 'yt_scrape_bot_check' });
}
if (/timed out|timeout|ETIMEDOUT|killed|SIGKILL/i.test(msg)) {
return Object.assign(new Error('YouTube scrape timeout'), { ytStatus: 504, code: 'yt_scrape_timeout' });
}
if (/unable to|unsupported|not available|private|deleted/i.test(msg)) {
return Object.assign(new Error(`YouTube scrape upstream: ${msg.slice(0, 160)}`), { ytStatus: 502, code: 'yt_scrape_upstream' });
}
return Object.assign(new Error(`YouTube scrape failed: ${msg.slice(0, 200)}`), { ytStatus: 502, code: 'yt_scrape_failed' });
}
async function runYtDlpFlat(queryOrUrl, { limit = 10, playlistStart = 1 } = {}) {
const bin = await resolveYtDlpBin();
const timeout = getYtDlpTimeoutMs();
const end = playlistStart + Math.max(1, Math.min(50, Number(limit || 10))) - 1;
const args = [
'--dump-single-json',
'--flat-playlist',
'--no-warnings',
'--no-check-certificates',
'--skip-download',
'--no-playlist',
'--playlist-start', String(playlistStart),
'--playlist-end', String(end),
...buildYtDlpExtraArgs(),
queryOrUrl,
];
try {
const { stdout } = await execFileAsync(bin, args, { timeout, maxBuffer: 16 * 1024 * 1024 });
return parseFlatPlaylistJson(stdout);
} catch (e) {
// yt-dlp peut écrire du JSON partiel sur stdout même en exit non-zero
const partial = e?.stdout ? parseFlatPlaylistJson(String(e.stdout)) : [];
if (partial.length) return partial;
throw classifyScrapeError(e);
}
}
/** Recherche sans clé. sort: relevance|date|views */
export async function searchViaScrape(q, opts = {}) {
const query = String(q || '').trim();
if (query.length < 2) return [];
const limit = Math.min(50, Math.max(1, Number(opts?.limit || 10)));
const page = Math.max(1, Number(opts?.page || 1));
const sort = String(opts?.sort || 'relevance').toLowerCase();
const perPage = limit;
const start = (page - 1) * perPage + 1;
let prefix = 'ytsearch';
if (sort === 'date') prefix = 'ytsearchdate';
// views : pas de préfixe natif stable -> ytsearch + tri local
const target = `${prefix}${perPage}:${query}`;
let items = await runYtDlpFlat(target, { limit: perPage, playlistStart: start });
if (sort === 'views') items = [...items].sort((a, b) => (b.views || 0) - (a.views || 0));
return items.slice(0, perPage);
}
const CHANNEL_TABS = { videos: 'videos', shorts: 'shorts', streams: 'streams', live: 'streams', playlists: 'playlists' };
/** Contenu chaîne sans clé. type: videos|shorts|playlists|live */
export async function fetchChannelViaScrape(externalId, { type = 'videos', page = 1, limit = 24 } = {}) {
const raw = String(externalId || '').trim().replace(/^@/, '');
if (!raw) return { items: [], nextPage: null };
const tab = CHANNEL_TABS[String(type)] || 'videos';
const perPage = Math.min(50, Math.max(1, Number(limit || 24)));
const pageNum = Math.max(1, Number(page || 1));
const start = (pageNum - 1) * perPage + 1;
// UC... -> /channel/UC.../tab, sinon -> /@/handle/tab
const base = /^UC[\w-]{20,}$/.test(raw)
? `https://www.youtube.com/channel/${raw}/${tab}`
: `https://www.youtube.com/@${raw}/${tab}`;
const items = await runYtDlpFlat(base, { limit: perPage, playlistStart: start });
// playlists renvoient des ids PL... : on les garde tels quels avec un shape playlist
if (tab === 'playlists') {
const mapped = items.map((s) => ({ id: s.id, title: s.title, thumbnail: s.thumbnail, videoCount: null, updatedAt: s.publishedAt || null }));
return { items: mapped, nextPage: items.length >= perPage ? pageNum + 1 : null };
}
let filtered = items;
if (type === 'shorts') filtered = items.filter((s) => !s.duration || s.duration <= 70);
return { items: filtered, nextPage: items.length >= perPage ? pageNum + 1 : null };
}
/** Trending sans clé. Les onglets /feed/trending ont été retirés côté YouTube
* (redirect home -> erreur tab) ; on tente les tabs puis repli ytsearch trié par vues. */
export async function getTrendingViaScrape(limit = 24) {
const n = Math.min(50, Math.max(1, Number(limit || 24)));
const tabs = [
'https://www.youtube.com/feed/trending',
'https://www.youtube.com/trending',
];
for (const url of tabs) {
try {
const items = await runYtDlpFlat(url, { limit: n, playlistStart: 1 });
if (items.length) return items;
} catch {}
}
// Repli : recherche générique triée par vues (0 quota, toujours disponible)
return runYtDlpFlat(`ytsearch${n}:top trending videos world`, { limit: n, playlistStart: 1 });
}
/** Résout un channel_id via scrape (imprime channel_id sans appel API). */
export async function resolveChannelIdViaScrape(externalId) {
const raw = String(externalId || '').trim();
if (/^UC[\w-]{20,}$/.test(raw)) return raw;
const handle = raw.replace(/^@/, '');
const bin = await resolveYtDlpBin().catch(() => null);
if (!bin) {
throw Object.assign(
new Error('yt-dlp introuvable (ni PATH ni bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
);
}
const timeout = Math.min(15000, getYtDlpTimeoutMs());
try {
const { stdout } = await execFileAsync(bin,
['--print', '%(channel_id)s', '--no-warnings', '--skip-download', '--playlist-end', '1', ...buildYtDlpExtraArgs(), `https://www.youtube.com/@${handle}/videos`],
{ timeout });
const id = String(stdout || '').trim().split('\n')[0].trim();
if (/^UC[\w-]{20,}$/.test(id)) return id;
} catch {}
return raw;
}
+270 -145
View File
@@ -9,6 +9,12 @@
* @property {string=} uploaderName
* @property {string=} type
*/
import {
getYouTubeKeys, isKeyFailure, getSearchMode, getScrapeTtlMs,
hashSearchKey, ytMetrics,
} from './youtube-common.mjs';
import { searchViaScrape } from './youtube-scrape.mjs';
import { searchViaInnerTube } from './youtube-innertube.mjs';
function parseISODurationToSeconds(iso) {
if (typeof iso !== 'string' || !iso) return 0;
@@ -20,50 +26,6 @@ function parseISODurationToSeconds(iso) {
return (hours * 3600) + (minutes * 60) + seconds;
}
/**
* Resolve configured YouTube API keys.
* Accepts YOUTUBE_API_KEYS as JSON array ('["k1","k2"]') or CSV ('k1,k2'),
* plus the legacy single YOUTUBE_API_KEY as fallback. De-duplicated.
* @returns {string[]}
*/
function getYouTubeKeys() {
const keys = [];
try {
const raw = process.env.YOUTUBE_API_KEYS;
if (raw && String(raw).trim() && String(raw).trim() !== 'undefined' && String(raw).trim() !== 'null') {
const s = String(raw).trim();
if (s.startsWith('[')) {
try {
const arr = JSON.parse(s);
if (Array.isArray(arr)) keys.push(...arr.map(v => String(v || '').trim()).filter(Boolean));
} catch {}
} else {
keys.push(...s.split(',').map(v => String(v || '').trim()).filter(Boolean));
}
}
} catch {}
try {
const single = process.env.YOUTUBE_API_KEY;
if (single && String(single).trim()) keys.push(String(single).trim());
} catch {}
return Array.from(new Set(keys.filter(Boolean)));
}
/**
* Classify a YouTube API failure as retryable with another key.
* - 400 API_KEY_INVALID / "API key expired" -> key is dead, try next.
* - 403 quotaExceeded / rateLimitExceeded / dailyLimitExceeded -> quota, try next.
*/
function isKeyFailure(status, data) {
try {
const reason = data?.error?.errors?.[0]?.reason || '';
const message = String(data?.error?.message || '');
if (status === 400 && (reason === 'API_KEY_INVALID' || /api key (expired|invalid)/i.test(message))) return true;
if (status === 403 && /quota|rateLimit|dailyLimit|userRateLimit/i.test(reason + ' ' + message)) return true;
} catch {}
return false;
}
/** GET JSON from YouTube, rotating through keys on key failures. */
async function ytFetchJson(base, paramsWithoutKey) {
const keys = getYouTubeKeys();
@@ -76,13 +38,19 @@ async function ytFetchJson(base, paramsWithoutKey) {
for (const key of keys) {
const params = new URLSearchParams(paramsWithoutKey);
params.set('key', key);
ytMetrics.apiCalls++;
try {
const { incYoutubeMetrics } = await import('../db.mjs').catch(() => ({}));
if (String(base).includes('/search')) incYoutubeMetrics?.({ apiCalls: 1, quotaUnits: 100 });
else incYoutubeMetrics?.({ apiCalls: 1, quotaUnits: 1 });
ytMetrics.quotaUnits += String(base).includes('/search') ? 100 : 1;
} catch {}
const resp = await fetch(`${base}?${params.toString()}`);
const data = await resp.json().catch(() => ({}));
if (resp.ok) return data;
lastStatus = resp.status;
lastData = data;
lastError = new Error(`YouTube API error: ${resp.status} ${data?.error?.message || ''}`.trim());
// Only rotate to the next key on key-attributable failures; otherwise fail fast.
if (!isKeyFailure(resp.status, data)) break;
console.warn(`[YouTube] key ...${String(key).slice(-4)} failed (${resp.status}), trying next key`);
}
@@ -92,121 +60,278 @@ async function ytFetchJson(base, paramsWithoutKey) {
throw err;
}
/**
* Parse a `suggestqueries.google.com` (client=youtube) response body into plain strings.
* The endpoint answers either bare JSON (`["q",["s1","s2"],...]`) or wrapped in
* `window.google.ac.h(...)`. Never throws: unparsable bodies yield [].
* @param {string} text
* @returns {string[]}
*/
export function parseYoutubeSuggestResponse(text) {
try {
const raw = String(text || '').trim();
if (!raw) return [];
let payload = raw;
const start = raw.indexOf('(');
const end = raw.lastIndexOf(')');
if (start !== -1 && end > start) payload = raw.slice(start + 1, end);
const parsed = JSON.parse(payload);
const candidates = Array.isArray(parsed) && Array.isArray(parsed[1]) ? parsed[1] : [];
return candidates
.map((s) => {
if (Array.isArray(s)) s = s[0];
return (typeof s === 'string' ? s : String(s ?? '')).trim();
})
.filter(Boolean)
.slice(0, 20);
} catch {
return [];
}
}
// Mémoire LRU process-wide pour le dispatcher (complète SQLite)
const memCache = new Map();
const MEM_MAX = Number(process.env.YT_SCRAPE_MEM_MAX || 300);
function memGet(k, ttl) {
const hit = memCache.get(k);
if (!hit) return null;
if ((Date.now() - hit.ts) >= ttl) { memCache.delete(k); return null; }
memCache.delete(k); memCache.set(k, hit);
return hit.items;
}
function memSet(k, items) {
if (memCache.has(k)) memCache.delete(k);
memCache.set(k, { ts: Date.now(), items });
while (memCache.size > MEM_MAX) { const o = memCache.keys().next().value; if (o === undefined) break; memCache.delete(o); }
}
export function ytScrapeCacheStats() {
return { memEntries: memCache.size, memMax: MEM_MAX };
}
async function searchViaApi(q, { limit = 10, page = 1, sort = 'relevance' } = {}) {
const keys = getYouTubeKeys();
if (!keys.length) {
throw Object.assign(new Error('YOUTUBE_API_KEY not configured'), { ytStatus: 503, code: 'youtube_api_key_unavailable' });
}
let order = 'relevance';
if (sort === 'date') order = 'date';
else if (sort === 'views') order = 'viewCount';
const perPage = Math.min(Math.max(1, Number(limit || 10)), 50);
const targetPage = Math.max(1, Number(page || 1));
let pageToken = '';
let currentPage = 1;
let lastItems = [];
while (currentPage <= targetPage) {
const params = {
part: 'snippet', q, type: 'video', maxResults: String(perPage), order,
videoEmbeddable: 'true', safeSearch: 'moderate',
};
if (pageToken) params.pageToken = pageToken;
const data = await ytFetchJson('https://www.googleapis.com/youtube/v3/search', params);
if (currentPage === targetPage) { lastItems = Array.isArray(data.items) ? data.items : []; break; }
const next = data.nextPageToken;
if (!next) { lastItems = []; break; }
pageToken = String(next);
currentPage++;
}
const videoIds = (lastItems || []).map((item) => item?.id?.videoId).filter(Boolean);
const detailsMap = new Map();
if (videoIds.length > 0) {
try {
const detailsData = await ytFetchJson('https://www.googleapis.com/youtube/v3/videos', {
part: 'contentDetails,statistics,status', id: videoIds.join(','),
});
for (const vid of detailsData?.items || []) if (vid?.id) detailsMap.set(vid.id, vid);
} catch (e) {
console.warn('[YouTube] details fetch failed, continuing without durations:', e?.message || e);
}
}
return (lastItems || []).map((item) => {
const videoId = item?.id?.videoId;
const snippet = item?.snippet || {};
const thumb = snippet.thumbnails?.high?.url || snippet.thumbnails?.medium?.url || snippet.thumbnails?.default?.url || undefined;
const details = videoId ? detailsMap.get(videoId) : null;
const duration = parseISODurationToSeconds(details?.contentDetails?.duration || '');
const views = details?.statistics?.viewCount != null ? Number(details.statistics.viewCount) : undefined;
const channelId = snippet.channelId || undefined;
const embeddable = details?.status ? details.status.embeddable !== false : undefined;
return {
title: snippet.title || '', id: videoId,
url: videoId ? `https://www.youtube.com/watch?v=${videoId}` : undefined,
thumbnail: thumb, uploaderName: snippet.channelTitle || undefined, type: 'video',
duration: duration > 0 ? duration : undefined, views,
publishedAt: snippet.publishedAt || undefined, channelId,
channelHandle: snippet.channelTitle || undefined, channelExternalId: channelId,
channelUrl: channelId ? `https://www.youtube.com/channel/${channelId}` : undefined, embeddable,
};
});
}
/** @type {{ id: 'yt', label: string, search: (q: string, opts: { limit: number, page?: number, sort?: 'relevance'|'date'|'views' }) => Promise<Suggestion[]> }} */
const handler = {
id: 'yt',
label: 'YouTube',
async search(q, opts) {
const { limit = 10, page = 1, sort = 'relevance' } = opts || {};
try {
const keys = getYouTubeKeys();
if (!keys.length) {
throw Object.assign(new Error('YOUTUBE_API_KEY not configured'), { ytStatus: 503, code: 'youtube_api_key_unavailable' });
}
let order = 'relevance';
if (sort === 'date') order = 'date';
else if (sort === 'views') order = 'viewCount';
// Iterate nextPageToken to reach the requested page (1-based)
const mode = getSearchMode();
const ttl = getScrapeTtlMs();
const perPage = Math.min(Math.max(1, Number(limit || 10)), 50);
const targetPage = Math.max(1, Number(page || 1));
let pageToken = '';
let currentPage = 1;
let lastItems = [];
while (currentPage <= targetPage) {
const params = {
part: 'snippet',
q: q,
type: 'video',
maxResults: String(perPage),
order,
// Ne retourner que des vidéos lisibles en embed (évite l'erreur 153 côté player).
// Les vidéos non-embeddables sont filtrées via le champ status ci-dessous pour /videos.
videoEmbeddable: 'true',
safeSearch: 'moderate',
};
if (pageToken) params.pageToken = pageToken;
const data = await ytFetchJson('https://www.googleapis.com/youtube/v3/search', params);
if (currentPage === targetPage) {
lastItems = Array.isArray(data.items) ? data.items : [];
break;
}
// Prepare for next iteration
const next = data.nextPageToken;
if (!next) {
// No more pages; requested page beyond available results
lastItems = [];
break;
}
pageToken = String(next);
currentPage++;
}
const videoIds = (lastItems || [])
.map(item => item?.id?.videoId)
.filter(Boolean);
const detailsMap = new Map();
if (videoIds.length > 0) {
const key = `yt|${hashSearchKey(`${String(q).toLowerCase().trim()}|${perPage}|${page}|${sort}|${mode}`)}`;
// 1) mémoire
const memHit = memGet(key, ttl);
if (memHit) { ytMetrics.scrapeHits++; return memHit; }
// 2) SQLite
try {
const detailsData = await ytFetchJson('https://www.googleapis.com/youtube/v3/videos', {
part: 'contentDetails,statistics,status',
id: videoIds.join(','),
});
for (const vid of detailsData?.items || []) {
if (vid?.id) detailsMap.set(vid.id, vid);
const { getCachedYoutubeSearch } = await import('../db.mjs');
const cached = getCachedYoutubeSearch(key);
if (cached?.items) {
ytMetrics.scrapeHits++;
memSet(key, cached.items);
return cached.items;
}
} catch (e) {
console.warn('[YouTube] details fetch failed, continuing without durations:', e?.message || e);
}
}
return (lastItems || []).map(item => {
const videoId = item?.id?.videoId;
const snippet = item?.snippet || {};
const thumb = snippet.thumbnails?.high?.url
|| snippet.thumbnails?.medium?.url
|| snippet.thumbnails?.default?.url
|| undefined;
const details = videoId ? detailsMap.get(videoId) : null;
const isoDuration = details?.contentDetails?.duration || '';
const duration = parseISODurationToSeconds(isoDuration);
const views = details?.statistics?.viewCount != null ? Number(details.statistics.viewCount) : undefined;
const channelId = snippet.channelId || undefined;
const channelHandle = snippet.channelTitle || undefined;
// status.embeddable === false -> le player renvoie l'erreur 153 ("Video configuration error").
const embeddable = details?.status ? details.status.embeddable !== false : undefined;
return {
title: snippet.title || '',
id: videoId,
url: videoId ? `https://www.youtube.com/watch?v=${videoId}` : undefined,
thumbnail: thumb,
uploaderName: snippet.channelTitle || undefined,
type: 'video',
duration: duration > 0 ? duration : undefined,
views,
publishedAt: snippet.publishedAt || undefined,
channelId,
channelHandle,
channelExternalId: channelId,
channelUrl: channelId ? `https://www.youtube.com/channel/${channelId}` : undefined,
embeddable,
} catch {}
const persist = async (items, source) => {
// Ne jamais persister un résultat vide : une page vide transitoire
// (continuation expirée, raté réseau partiel) ne doit pas empoisonner
// le cache et bloquer les requêtes suivantes.
if (!Array.isArray(items) || items.length === 0) return items;
memSet(key, items);
try {
const { setCachedYoutubeSearch, incYoutubeMetrics } = await import('../db.mjs');
setCachedYoutubeSearch(key, q, items, source, ttl);
if (source === 'scrape') incYoutubeMetrics({ scrapeCalls: 1 });
} catch {}
return items;
};
});
const t0 = Date.now();
const tryScrape = async () => {
ytMetrics.scrapeCalls++;
try { const { incYoutubeMetrics } = await import('../db.mjs'); incYoutubeMetrics({ scrapeCalls: 1 }); } catch {}
return searchViaScrape(q, { limit: perPage, page, sort });
};
const tryApi = () => searchViaApi(q, { limit: perPage, page, sort });
const tryInnerTube = () => searchViaInnerTube(q, { limit: perPage, page, sort });
const log = (source, extra = '') => console.log(`[YT search] source=${source} mode=${mode} latency=${Date.now() - t0}ms results=${extra}`);
try {
if (mode === 'api-only') {
const items = await tryApi();
log('api', items.length);
return persist(items, 'api');
}
if (mode === 'scrape-only') {
const items = await tryScrape();
log('scrape', items.length);
return persist(items, 'scrape');
}
if (mode === 'innertube-only') {
const items = await tryInnerTube();
log('innertube', items.length);
return persist(items, 'innertube');
}
if (mode === 'api-first') {
try {
const items = await tryApi();
log('api', items.length);
return persist(items, 'api');
} catch (apiErr) {
ytMetrics.fallbacks++;
console.warn(`[YT search] api-first fallback to scrape: ${apiErr?.message || apiErr}`);
const items = await tryScrape();
log('scrape(fallback)', items.length);
return persist(items, 'scrape');
}
}
if (mode === 'scrape-first') {
// scrape -> API (comportement historique, sans InnerTube)
try {
const items = await tryScrape();
log('scrape', items.length);
return persist(items, 'scrape');
} catch (scrapeErr) {
ytMetrics.scrapeErrors++;
ytMetrics.fallbacks++;
console.warn(`[YT search] scrape failed (${scrapeErr?.code || 'unknown'}), fallback to api`);
try {
const items = await tryApi();
log('api(fallback)', items.length);
return persist(items, 'api');
} catch (apiErr) {
const noBinary = scrapeErr?.code === 'yt_scrape_no_binary';
const noKey = apiErr?.code === 'youtube_api_key_unavailable';
if (noBinary && noKey) {
const err = new Error('YouTube indisponible : yt-dlp introuvable (définissez YT_DLP_PATH) et aucune clé YOUTUBE_API_KEY configurée.');
err.ytStatus = 503; err.code = 'youtube_no_source';
throw err;
}
console.error(`[YT search] both sources failed. scrape=${scrapeErr?.message} api=${apiErr?.message}`);
throw scrapeErr;
}
}
}
// innertube-first (défaut) : InnerTube -> scrape yt-dlp -> API officielle.
// Chaque couche ne fait jamais échouer la recherche à elle seule.
const errors = [];
try {
const items = await tryInnerTube();
log('innertube', items.length);
return persist(items, 'innertube');
} catch (itErr) {
ytMetrics.innertubeErrors = (ytMetrics.innertubeErrors || 0) + 1;
ytMetrics.fallbacks++;
errors.push(`innertube=${itErr?.message}`);
console.warn(`[YT search] innertube failed (${itErr?.code || 'unknown'}), fallback to scrape`);
}
try {
const items = await tryScrape();
log('scrape(fallback)', items.length);
return persist(items, 'scrape');
} catch (scrapeErr) {
ytMetrics.scrapeErrors++;
ytMetrics.fallbacks++;
errors.push(`scrape=${scrapeErr?.message}`);
console.warn(`[YT search] scrape failed (${scrapeErr?.code || 'unknown'}), fallback to api`);
}
try {
const items = await tryApi();
log('api(fallback)', items.length);
return persist(items, 'api');
} catch (apiErr) {
errors.push(`api=${apiErr?.message}`);
const noSource = !getYouTubeKeys().length;
if (noSource) {
const err = new Error(`YouTube indisponible (${errors.join(' | ').slice(0, 220)}). Vérifiez le réseau/YT_DLP_PATH ou ajoutez YOUTUBE_API_KEYS.`);
err.ytStatus = 503; err.code = 'youtube_no_source';
throw err;
}
console.error(`[YT search] all sources failed. ${errors.join(' | ')}`);
const err = new Error(errors[0] || 'search_failed');
err.ytStatus = 502; err.code = 'youtube_all_sources_failed';
throw err;
}
} catch (error) {
// Log explicite (clé expirée / quota) puis remonte l'erreur pour que
// /api/search la reporte dans `errors.yt` au lieu d'un groupe vide silencieux.
const status = error?.ytStatus || 500;
console.error(`YouTube search error (status ${status}):`, error?.message || error);
throw error;
}
},
async suggest(q, opts) {
const limit = Math.min(20, Math.max(1, Number(opts?.limit || 10)));
const query = String(q || '').trim();
if (query.length < 2) return [];
try {
const params = new URLSearchParams({ client: 'youtube', ds: 'yt', hl: 'fr', q: query });
const headers = { 'User-Agent': 'Mozilla/5.0' };
const proxy = String(process.env.YT_EGRESS_PROXY || '').trim();
const fetchOpts = { headers, signal: AbortSignal.timeout(6000) };
// Note : suggest reste sans proxy (endpoint Google public léger) sauf si egress configuré côté infra.
if (proxy) console.debug?.('[YT suggest] egress proxy configured, direct fetch kept for suggest');
const resp = await fetch(`https://suggestqueries.google.com/complete/search?${params.toString()}`, fetchOpts);
if (!resp.ok) return [];
const text = await resp.text();
return parseYoutubeSuggestResponse(text).slice(0, limit);
} catch {
return [];
}
}
};
+130
View File
@@ -0,0 +1,130 @@
// Step 15 — Suggest tests: unit parsing/dedup (offline) + API contract (isolated server).
// Run with: npm run test:suggest
// Scenario coverage (todo Step 15):
// - providers=yt returns only yt suggestions
// - unknown provider falls back to the full registry
// Assertions are on fan-out shape, never on external provider content.
import fs from 'node:fs';
import path from 'node:path';
import os from 'node:os';
import { spawn } from 'node:child_process';
import net from 'node:net';
function expect(cond, msg) {
if (!cond) throw new Error(`Assertion failed: ${msg}`);
}
function logOk(msg) { console.log(`✓ ${msg}`); }
// ---------- 1) Offline unit tests: parsers + graceful degradation ----------
const yt = (await import('../providers/youtube.mjs')).default;
const { parseYoutubeSuggestResponse } = await import('../providers/youtube.mjs');
const dm = (await import('../providers/dailymotion.mjs')).default;
// Bare JSON payload
expect(
JSON.stringify(parseYoutubeSuggestResponse('["tut",["tutoriel angular","tutoriel android"],[],{}]')) ===
JSON.stringify(['tutoriel angular', 'tutoriel android']),
'bare JSON suggest payload parses',
);
logOk('parseYoutubeSuggestResponse bare JSON');
// window.google.ac.h(...) wrapped payload
expect(
JSON.stringify(
parseYoutubeSuggestResponse('window.google.ac.h(["tut",["tutoriel angular"],[],{}])'),
) === JSON.stringify(['tutoriel angular']),
'wrapped JSONP suggest payload parses',
);
logOk('parseYoutubeSuggestResponse wrapped payload');
// Array-wrapped suggestions (newer format: [text, type, ...])
expect(
JSON.stringify(parseYoutubeSuggestResponse('["tut",[["tutoriel angular",0],["tutoriel android",0]]]')) ===
JSON.stringify(['tutoriel angular', 'tutoriel android']),
'array-wrapped suggestions keep text only',
);
logOk('parseYoutubeSuggestResponse array-wrapped entries');
// Garbage / empty payloads never throw
expect(parseYoutubeSuggestResponse('not json at all').length === 0, 'garbage payload yields []');
expect(parseYoutubeSuggestResponse('').length === 0, 'empty payload yields []');
logOk('parseYoutubeSuggestResponse graceful degradation');
// min-length guard: no network below 2 chars (would throw offline otherwise)
expect(JSON.stringify(await yt.suggest('a')) === '[]', 'yt.suggest short query returns [] without network');
expect(JSON.stringify(await dm.suggest('x')) === '[]', 'dm.suggest short query returns [] without network');
logOk('suggest min-length guard (no network)');
// ---------- 2) API contract against a real isolated server ----------
const PORT = await new Promise((resolve) => {
const srv = net.createServer();
srv.listen(0, '127.0.0.1', () => {
const p = srv.address().port;
srv.close(() => resolve(p));
});
});
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'newtube-suggest-test-'));
const server = spawn(process.execPath, ['./server/index.mjs'], {
env: { ...process.env, PORT: String(PORT), NEWTUBE_DB_FILE: path.join(tmpDir, 'suggest.db'), JWT_SECRET: 'suggest-test-secret', NODE_ENV: 'test' },
stdio: ['ignore', 'pipe', 'pipe'],
cwd: path.resolve(import.meta.dirname, '..', '..'),
});
let serverLogs = '';
server.stdout.on('data', (d) => { serverLogs += d.toString(); });
server.stderr.on('data', (d) => { serverLogs += d.toString(); });
const baseUrl = `http://127.0.0.1:${PORT}`;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function waitForServer(timeoutMs = 20000) {
const start = Date.now();
while (Date.now() - start < timeoutMs) {
try {
const res = await fetch(`${baseUrl}/`);
if (res.status < 500) return true;
} catch {}
await sleep(300);
}
return false;
}
let failures = 0;
async function scenario(name, fn) {
try { await fn(); logOk(name); }
catch (e) { failures += 1; console.error(`✗ ${name}`); console.error(e?.stack || String(e)); }
}
try {
const up = await waitForServer();
expect(up, 'isolated server starts');
await scenario('q too short -> 400', async () => {
const res = await fetch(`${baseUrl}/api/search/suggest?q=a`);
expect(res.status === 400, `expected 400, got ${res.status}`);
});
await scenario('providers=yt returns only yt suggestions', async () => {
const res = await fetch(`${baseUrl}/api/search/suggest?q=tutoriel&providers=yt&limit=5`);
expect(res.status === 200, `expected 200, got ${res.status}`);
const body = await res.json();
expect(body.q === 'tutoriel', 'q echoed');
expect(JSON.stringify(Object.keys(body.groups)) === JSON.stringify(['yt']), `only yt group (got ${Object.keys(body.groups)})`);
expect(Array.isArray(body.groups.yt), 'yt group is an array');
expect(body.groups.yt.length <= 5, 'limit respected');
});
await scenario('unknown provider falls back to the full registry', async () => {
const res = await fetch(`${baseUrl}/api/search/suggest?q=tutoriel&providers=xx,yy&limit=3`);
expect(res.status === 200, `expected 200, got ${res.status}`);
const body = await res.json();
const keys = Object.keys(body.groups).sort();
expect(JSON.stringify(keys) === JSON.stringify(['dm', 'od', 'pt', 'ru', 'tw', 'yt']), `full registry fallback (got ${keys})`);
});
} finally {
server.kill();
}
if (failures > 0) {
console.error(`\n${failures} suggest scenario(s) failed. Server logs:\n${serverLogs}`);
process.exit(1);
}
console.log('\nAll suggest tests passed.');
+212
View File
@@ -0,0 +1,212 @@
// Step 16 — Transcript tests (offline fixtures for parsers + API contract).
// Run with: npm run test:transcript
import { describe, it } from 'node:test';
import assert from 'node:assert/strict';
import fs from 'node:fs';
import path from 'node:path';
import os from 'node:os';
import { spawn } from 'node:child_process';
import net from 'node:net';
import {
pickTrack,
parseJson3,
parseVtt,
parseTrackText,
parseXmlCaptions,
orderedTracks,
looksLikeHtmlError,
capLines,
normalizeTranscriptProvider,
} from '../transcript.mjs';
const track = (url, ext = 'vtt') => ({ url, ext });
describe('pickTrack', () => {
it('prefers manual subtitles over automatic captions', () => {
const json = {
subtitles: { en: [track('https://x/en.vtt')] },
automatic_captions: { en: [track('https://x/en.auto.vtt')] },
};
const { track: t, lang } = pickTrack(json, 'en');
assert.equal(t.url, 'https://x/en.vtt');
assert.equal(lang, 'en');
});
it('falls back fr -> fr-* -> auto -> first available', () => {
const json = {
subtitles: { 'fr-CA': [track('https://x/frca.vtt')] },
automatic_captions: { en: [track('https://x/en.auto.vtt')] },
};
assert.equal(pickTrack(json, 'fr').lang, 'fr-CA');
// 'de' matches nothing: first available (manual preferred)
const fb = pickTrack(json, 'de');
assert.equal(fb.lang, 'fr-CA');
// auto-only dict
const autoOnly = pickTrack({ automatic_captions: { en: [track('https://x/e.vtt')] } }, 'en');
assert.equal(autoOnly.track.url, 'https://x/e.vtt');
});
it('returns null track when no subtitles exist', () => {
assert.deepEqual(pickTrack({}, 'fr'), { track: null, languages: [], lang: null });
assert.deepEqual(pickTrack({ subtitles: {}, automatic_captions: {} }, 'en'), {
track: null,
languages: [],
lang: null,
});
});
it('exposes the language list for the UI selector', () => {
const { languages } = pickTrack(
{ subtitles: { fr: [track('a')], en: [track('b')] }, automatic_captions: { es: [track('c')] } },
'fr',
);
assert.deepEqual([...languages].sort(), ['en', 'es', 'fr']);
});
});
describe('parseJson3', () => {
it('converts events with segs to normalized lines', () => {
const lines = parseJson3({
events: [
{ tStartMs: 0, dDurationMs: 2500, segs: [{ utf8: 'Bonjour' }] },
{ tStartMs: 2500, dDurationMs: 3100, segs: [{ utf8: 'Bien' }, { utf8: 'venue' }] },
],
});
assert.deepEqual(lines, [
{ t: 0, dur: 2.5, text: 'Bonjour' },
{ t: 2.5, dur: 3.1, text: 'Bienvenue' },
]);
});
it('drops events without segs or with empty text', () => {
const lines = parseJson3({ events: [{ tStartMs: 0 }, { tStartMs: 1, dDurationMs: 1, segs: [{ utf8: ' ' }] }] });
assert.deepEqual(lines, []);
assert.deepEqual(parseJson3({}), []);
});
});
describe('parseVtt', () => {
const vtt = `WEBVTT
00:00:00.000 --> 00:00:02.500
Bonjour <c.colorE5E5E5>à tous</c>
12
00:00:02.500 --> 00:00:05.600
Bienvenue
dans cette vidéo
`;
it('parses cues with identifiers, tags and multi-line content', () => {
const lines = parseVtt(vtt);
assert.equal(lines.length, 2);
assert.equal(lines[0].t, 0);
assert.equal(lines[0].dur, 2.5);
assert.equal(lines[0].text, 'Bonjour à tous');
assert.equal(lines[1].text, 'Bienvenue dans cette vidéo');
assert.ok(Math.abs(lines[1].t - 2.5) < 1e-9);
});
it('supports MM:SS.mmm timestamps and decodes entities', () => {
const lines = parseVtt('WEBVTT\n\n01:02.000 --> 01:04.500\nFish &amp; chips\n');
assert.equal(lines.length, 1);
assert.equal(lines[0].t, 62);
assert.equal(lines[0].text, 'Fish & chips');
});
it('ignores garbage blocks', () => {
assert.deepEqual(parseVtt('WEBVTT\n\nNOTE nothing here\n'), []);
});
});
describe('orderedTracks + parseTrackText + XML captions', () => {
const auto = {
fr: [{ url: 'https://x/fr.json3', ext: 'json3' }],
en: [{ url: 'https://x/en.json3', ext: 'json3' }],
es: [{ url: 'https://x/es.vtt', ext: 'vtt' }],
};
it('tries requested lang first, then original en', () => {
const langs = orderedTracks({ subtitles: {}, automatic_captions: auto }, 'fr').map((o) => o.lang);
assert.deepEqual(langs, ['fr', 'en', 'es']);
});
it('dedupes identical track URLs', () => {
const dup = { subtitles: {}, automatic_captions: { en: [{ url: 'https://x/same', ext: 'vtt' }], fr: [{ url: 'https://x/same', ext: 'vtt' }] } };
assert.equal(orderedTracks(dup, 'fr').length, 1);
});
it('detects HTML error pages instead of JSON-parsing them', () => {
assert.equal(looksLikeHtmlError('<!DOCTYPE html><html>Sorry</html>'), true);
assert.equal(looksLikeHtmlError('WEBVTT\n\n00:00:00.000 --> 1'), false);
assert.deepEqual(parseTrackText('<!DOCTYPE html> nope', 'json3'), []);
});
it('parses srv/ttml XML captions', () => {
const srv = parseXmlCaptions('<transcript><p t="0" d="2500">Hello</p></transcript>');
assert.equal(srv.length, 1);
assert.equal(srv[0].text, 'Hello');
const ttml = parseXmlCaptions('<tt><body><div><p begin="00:00:01.000" end="00:00:03.500">Bonjour</p></div></body></tt>');
assert.equal(ttml.length, 1);
assert.equal(ttml[0].t, 1);
});
it('parseTrackText handles json3 and vtt payloads', () => {
const j = parseTrackText(JSON.stringify({ events: [{ tStartMs: 0, dDurationMs: 1000, segs: [{ utf8: 'Hi' }] }] }), 'json3');
assert.equal(j.length, 1);
const v = parseTrackText('WEBVTT\n\n00:00:00.000 --> 00:00:01.000\nHello\n', 'vtt');
assert.equal(v.length, 1);
});
});
describe('capLines + normalizeTranscriptProvider', () => { it('truncates very long transcripts', () => {
const lines = Array.from({ length: 10 }, (_, i) => ({ t: i, dur: 1, text: `l${i}` }));
assert.equal(capLines(lines, 3).length, 3);
});
it('maps short and long provider ids', () => {
assert.equal(normalizeTranscriptProvider('yt'), 'youtube');
assert.equal(normalizeTranscriptProvider('youtube'), 'youtube');
assert.equal(normalizeTranscriptProvider('dm'), 'dailymotion');
assert.equal(normalizeTranscriptProvider('pt'), 'peertube');
assert.equal(normalizeTranscriptProvider('xx'), null);
});
});
describe('API contract', () => {
it('rejects unknown providers with 400 { available: false }', async () => {
const port = await new Promise((resolve) => {
const srv = net.createServer();
srv.listen(0, '127.0.0.1', () => {
const p = srv.address().port;
srv.close(() => resolve(p));
});
});
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'newtube-transcript-test-'));
const server = spawn(process.execPath, ['./server/index.mjs'], {
env: {
...process.env,
PORT: String(port),
NEWTUBE_DB_FILE: path.join(tmpDir, 'transcript.db'),
JWT_SECRET: 'transcript-test-secret',
NODE_ENV: 'test',
},
stdio: ['ignore', 'pipe', 'pipe'],
cwd: path.resolve(import.meta.dirname, '..', '..'),
});
const baseUrl = `http://127.0.0.1:${port}`;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
try {
let up = false;
const start = Date.now();
while (Date.now() - start < 20000) {
try {
const res = await fetch(`${baseUrl}/`);
if (res.status < 500) { up = true; break; }
} catch {}
await sleep(300);
}
assert.ok(up, 'isolated server starts');
const res = await fetch(`${baseUrl}/api/transcript/xx/abc123?lang=fr`);
assert.equal(res.status, 400);
const body = await res.json();
assert.equal(body.available, false);
assert.ok(typeof body.error === 'string');
} finally {
server.kill();
}
});
});
+131
View File
@@ -0,0 +1,131 @@
// Step 18 — tests offline InnerTube (AUCUN réseau : mappers purs + modes).
import assert from 'node:assert/strict';
import { mapVideoNode, mapLockupView, mapNodes, mapCaptionTracks, parseViewsText, parseDurationLabel } from '../providers/youtube-innertube.mjs';
import { pickTrack, orderedTracks } from '../transcript.mjs';
import { getSearchMode, YT_SEARCH_MODES } from '../providers/youtube-common.mjs';
function ok(msg) { console.log(`✓ ${msg}`); }
// 1) parseViewsText FR/EN
{
assert.equal(parseViewsText('81 973 vues'), 81973);
assert.equal(parseViewsText('1,2 M vues'), 1200000);
assert.equal(parseViewsText('57K'), 57000);
assert.equal(parseViewsText('3 k vues'), 3000);
assert.equal(parseViewsText(''), undefined);
assert.equal(parseViewsText(null), undefined);
ok('parseViewsText FR/EN');
}
// 1b) parseDurationLabel (a11y)
{
assert.equal(parseDurationLabel("What's new in Angular 18 2 minutes, 27 seconds"), 147);
assert.equal(parseDurationLabel('Live 1 heure, 5 minutes'), 3900);
assert.equal(parseDurationLabel('nonsense'), undefined);
ok('parseDurationLabel');
}
// 2) mapVideoNode type Video (forme réelle youtubei.js v18)
{
const m = mapVideoNode({
type: 'Video', id: 'abc123',
title: { text: 'Demo' },
duration: { text: '15:22', seconds: 922 },
thumbnails: [{ url: 'http://t/low.jpg' }, { url: 'http://t/high.jpg' }],
author: { id: 'UC999', name: 'Chaine' },
published: { text: 'il y a 2 ans' },
view_count: { text: '81 973 vues' },
badges: [{ label: 'Sous-titres' }],
is_live: false,
});
assert.equal(m.id, 'abc123');
assert.equal(m.url, 'https://www.youtube.com/watch?v=abc123');
assert.equal(m.thumbnail, 'http://t/high.jpg');
assert.equal(m.duration, 922);
assert.equal(m.views, 81973);
assert.equal(m.channelId, 'UC999');
assert.deepEqual(m.badges, ['Sous-titres']);
assert.equal(m.isShort, undefined);
ok('mapVideoNode Video');
}
// 3) Shorts/Reel -> isShort ; live -> isLive ; bruit filtré
{
const s = mapVideoNode({ type: 'ShortsLockupView', id: 'sh1', title: { text: 'Short' }, thumbnails: [{ url: 'u' }] });
assert.equal(s.isShort, true);
const l = mapVideoNode({ type: 'Video', id: 'lv1', title: { text: 'Live' }, thumbnails: [{ url: 'u' }], is_live: true });
assert.equal(l.isLive, true);
assert.equal(mapVideoNode({ type: 'Playlist', id: 'PL1', title: { text: 'PL' } }), null);
assert.equal(mapVideoNode({ type: 'Channel', id: 'UC1', title: { text: 'C' } }), null);
assert.equal(mapVideoNode(null), null);
assert.equal(mapVideoNode({ type: 'Video', title: { text: 'no id' } }), null);
ok('mapVideoNode shorts/live/filtrage');
}
// 4) mapNodes déduplique
{
const mk = (id) => ({ type: 'Video', id, title: { text: id }, thumbnails: [{ url: 'u' }] });
const out = mapNodes([mk('a'), mk('a'), mk('b'), null]);
assert.deepEqual(out.map((x) => x.id), ['a', 'b']);
ok('mapNodes dédup');
}
// 4b) mapLockupView (renderer moderne du watch-next façon SmartTube)
{
const m = mapLockupView({
type: 'LockupView', content_id: '0wIL6d5TxGc', content_type: 'VIDEO',
content_image: { image: [{ url: 'http://t/small.jpg', width: 168 }, { url: 'http://t/big.jpg', width: 336 }] },
metadata: {
title: { text: "What's new" },
metadata: { metadata_rows: [
{ metadata_parts: [{ text: { text: 'Rainer Hahnekamp' } }] },
{ metadata_parts: [{ text: { text: '57K' } }] },
] },
image: { a11y_label: 'Go to channel Rainer Hahnekamp' },
},
renderer_context: { accessibility_context: { label: "What's new 2 minutes, 27 seconds" } },
});
assert.equal(m.id, '0wIL6d5TxGc');
assert.equal(m.url, 'https://www.youtube.com/watch?v=0wIL6d5TxGc');
assert.equal(m.thumbnail, 'http://t/big.jpg');
assert.equal(m.uploaderName, 'Rainer Hahnekamp');
assert.equal(m.views, 57000);
assert.equal(m.duration, 147);
assert.equal(mapVideoNode({ type: 'LockupView', content_id: 'x', content_type: 'VIDEO', metadata: { title: { text: 'T' } } }).id, 'x');
assert.equal(mapLockupView({ type: 'LockupView', content_type: 'PLAYLIST', content_id: 'PL1', metadata: { title: { text: 'PL' } } }), null);
ok('mapLockupView watch-next');
}
// 5) modes innertube connus
{
assert.ok(YT_SEARCH_MODES.includes('innertube-first'));
assert.ok(YT_SEARCH_MODES.includes('innertube-only'));
const prev = process.env.YT_SEARCH_MODE;
process.env.YT_SEARCH_MODE = 'innertube-only';
assert.equal(getSearchMode(), 'innertube-only');
if (prev === undefined) delete process.env.YT_SEARCH_MODE; else process.env.YT_SEARCH_MODE = prev;
ok('modes innertube-first/only');
}
// 6) mapCaptionTracks : forme InnerTube -> pseudo dump yt-dlp, interop pickTrack
{
const cap = mapCaptionTracks([
{ base_url: 'https://x/t?caps=asr&x=1', language_code: 'en', kind: 'asr', name: { text: 'English (auto-generated)' } },
{ base_url: 'https://x/t?caps=m&x=2', language_code: 'fr', name: { text: 'Français' } },
{ base_url: '', language_code: 'de', name: 'Deutsch' },
]);
assert.deepEqual(cap.languages.sort(), ['en', 'fr']);
assert.equal(cap.trackCount, 2);
assert.match(cap.automatic_captions.en[0].url, /fmt=json3/);
assert.equal(cap.automatic_captions.en[0].ext, 'json3');
assert.equal(cap.subtitles.fr[0].name, 'Français');
// Interop : le pseudo dump passe dans pickTrack/orderedTracks sans modification
const picked = pickTrack({ subtitles: cap.subtitles, automatic_captions: cap.automatic_captions }, 'fr');
assert.equal(picked.lang, 'fr');
assert.ok(picked.track.url.includes('fmt=json3'));
const ordered = orderedTracks({ subtitles: cap.subtitles, automatic_captions: cap.automatic_captions }, 'de');
assert.ok(ordered.length >= 2);
ok('mapCaptionTracks + interop pickTrack/orderedTracks');
}
console.log('youtube-innertube tests passed');
+98
View File
@@ -0,0 +1,98 @@
// Step 17 — tests offline : parsers scrape + dispatcher (AUCUN réseau, AUCUNE clé).
import assert from 'node:assert/strict';
import { mapFlatEntry, parseFlatPlaylistJson, classifyScrapeError } from '../providers/youtube-scrape.mjs';
import { getSearchMode, hashSearchKey, buildYtDlpExtraArgs, resolveYtDlpBin, resetYtDlpBinCache } from '../providers/youtube-common.mjs';
function ok(msg) { console.log(`✓ ${msg}`); }
// 1) mapFlatEntry
{
const s = mapFlatEntry({
id: 'dQw4w9WgXcQ', title: 'Demo', channel_id: 'UC123', channel: 'Chaine',
duration: 212, view_count: 1000, timestamp: 1700000000,
thumbnails: [{ url: 'http://t/low.jpg', height: 90 }, { url: 'http://t/high.jpg', height: 720 }],
});
assert.equal(s.id, 'dQw4w9WgXcQ');
assert.equal(s.url, 'https://www.youtube.com/watch?v=dQw4w9WgXcQ');
assert.equal(s.thumbnail, 'http://t/high.jpg');
assert.equal(s.duration, 212);
ok('mapFlatEntry maps flat entry to Suggestion');
}
assert.equal(mapFlatEntry(null), null);
assert.equal(mapFlatEntry({}), null);
ok('mapFlatEntry null-safe');
// 2) parseFlatPlaylistJson : objet + NDJSON + vide
{
const raw = JSON.stringify({ entries: [{ id: 'a', title: 'A' }, { id: '', title: 'x' }, null] });
assert.equal(parseFlatPlaylistJson(raw).length, 1);
ok('parseFlatPlaylistJson objet .entries');
}
{
const raw = `{"id":"a","title":"A"}\n{"id":"b","title":"B"}\nnot-json\n`;
assert.equal(parseFlatPlaylistJson(raw).length, 2);
ok('parseFlatPlaylistJson NDJSON tolerant');
}
assert.deepEqual(parseFlatPlaylistJson(''), []);
assert.deepEqual(parseFlatPlaylistJson('{{{'), []);
ok('parseFlatPlaylistJson vide/invalide -> []');
// 3) getSearchMode défaut innertube-first + normalisation
{
const prev = process.env.YT_SEARCH_MODE;
delete process.env.YT_SEARCH_MODE;
assert.equal(getSearchMode(), 'innertube-first');
process.env.YT_SEARCH_MODE = 'API-FIRST';
assert.equal(getSearchMode(), 'api-first');
process.env.YT_SEARCH_MODE = 'scrape-first';
assert.equal(getSearchMode(), 'scrape-first');
process.env.YT_SEARCH_MODE = 'nonsense';
assert.equal(getSearchMode(), 'innertube-first');
if (prev === undefined) delete process.env.YT_SEARCH_MODE; else process.env.YT_SEARCH_MODE = prev;
ok('getSearchMode défaut + normalisation');
}
// 4) hash stable + extra args jamais de secret en clair autre que proxy path
{
assert.equal(hashSearchKey('a'), hashSearchKey('a'));
assert.notEqual(hashSearchKey('a'), hashSearchKey('b'));
const args = buildYtDlpExtraArgs();
assert.ok(Array.isArray(args));
ok('hashSearchKey + buildYtDlpExtraArgs');
}
// 5) dispatcher scrape-only sans clé : ne doit PAS exiger YOUTUBE_API_KEY au import,
// et searchViaScrape vide sur query courte
{
const { searchViaScrape } = await import('../providers/youtube-scrape.mjs');
assert.deepEqual(await searchViaScrape('a'), []);
ok('searchViaScrape query courte -> [] sans réseau');
}
// 6) classifyScrapeError : ENOENT -> yt_scrape_no_binary (message actionnable, jamais de spawn brut en UI)
{
const e = classifyScrapeError(Object.assign(new Error('spawn yt-dlp ENOENT'), { code: 'ENOENT', errno: 'ENOENT' }));
assert.equal(e.code, 'yt_scrape_no_binary');
assert.equal(e.ytStatus, 503);
assert.match(e.message, /YT_DLP_PATH/);
const passthrough = classifyScrapeError(Object.assign(new Error('x'), { code: 'yt_scrape_no_binary' }));
assert.equal(passthrough.code, 'yt_scrape_no_binary');
ok('classifyScrapeError ENOENT -> yt_scrape_no_binary');
}
// 7) resolveYtDlpBin : YT_DLP_PATH inexistant -> fallback PATH/bundled (jamais de throw si un binaire existe)
{
resetYtDlpBinCache();
const prev = process.env.YT_DLP_PATH;
process.env.YT_DLP_PATH = '/nonexistent/yt-dlp';
try {
const bin = await resolveYtDlpBin();
assert.ok(typeof bin === 'string' && bin.length > 0);
ok(`resolveYtDlpBin fallback -> ${bin}`);
} finally {
resetYtDlpBinCache();
if (prev === undefined) delete process.env.YT_DLP_PATH; else process.env.YT_DLP_PATH = prev;
}
}
console.log('youtube-scrape tests passed');
+309
View File
@@ -0,0 +1,309 @@
// Step 16 — Pure transcript helpers (no network, no DB).
// Data source: `yt-dlp --dump-single-json --skip-download` exposes
// `subtitles` (manual) and `automatic_captions` (auto-generated).
// Track entries are arrays of { url, ext, name } (or single objects).
/**
* @typedef {{ t: number, dur: number, text: string }} TranscriptLine
*/
const MAX_LINES_DEFAULT = Number(process.env.TRANSCRIPT_MAX_LINES || 5000);
/** Normalize a language code for comparison (lowercase, `_` -> `-`). */
function normLang(code) {
return String(code || '').trim().toLowerCase().replace(/_/g, '-');
}
/** Pick the first usable track object from an array-or-single entry. */
function firstTrack(entry) {
if (!entry) return null;
const list = Array.isArray(entry) ? entry : [entry];
return list.find((t) => t && typeof t.url === 'string' && t.url) || null;
}
function trackExtOf(track) {
const ext = String(track?.ext || '').toLowerCase();
if (ext) return ext;
try {
const u = new URL(String(track?.url || ''));
const m = /\.([a-z0-9]+)(?:[?#]|$)/i.exec(u.pathname || '');
if (m) return m[1].toLowerCase();
} catch {}
return '';
}
/** Find a language key in `dict` matching `code` exactly or by prefix (`fr` -> `fr-*`). */
function findLangKey(dict, code) {
const want = normLang(code);
if (!want || !dict) return null;
const keys = Object.keys(dict);
const exact = keys.find((k) => normLang(k) === want);
if (exact) return exact;
// `fr-ca` requested, `fr` available (and vice versa): match on primary subtag
const primary = want.split('-')[0];
return keys.find((k) => normLang(k).split('-')[0] === primary) || null;
}
/**
* Select the best subtitle track.
* Priority: manual exact > manual prefix/primary > auto exact > auto prefix/primary > first available.
* @param {any} json yt-dlp dump-single-json
* @param {string} [lang]
* @returns {{ track: any|null, languages: string[], lang: string|null }}
*/
export function pickTrack(json, lang = 'fr') {
const manual = (json && json.subtitles) || {};
const auto = (json && json.automatic_captions) || {};
const manualKeys = Object.keys(manual);
const autoKeys = Object.keys(auto);
const languages = Array.from(new Set([...manualKeys, ...autoKeys]));
if (languages.length === 0) return { track: null, languages: [], lang: null };
const manualKey = findLangKey(manual, lang);
if (manualKey) {
const track = firstTrack(manual[manualKey]);
if (track) return { track, languages, lang: manualKey };
}
const autoKey = findLangKey(auto, lang);
if (autoKey) {
const track = firstTrack(auto[autoKey]);
if (track) return { track, languages, lang: autoKey };
}
// Fallback: first available track (manual preferred)
for (const key of manualKeys) {
const track = firstTrack(manual[key]);
if (track) return { track, languages, lang: key };
}
for (const key of autoKeys) {
const track = firstTrack(auto[key]);
if (track) return { track, languages, lang: key };
}
return { track: null, languages, lang: null };
}
/**
* Order candidate tracks to try when the preferred one cannot be fetched.
* YouTube auto-translated tracks (`tlang=xx`) are often rate-limited (HTTP 429
* or "Sorry" HTML pages) while the original language still works: always try
* the requested language first, then the original (`en`), then everything
* else (manual preferred). De-duplicated by URL.
* @param {any} json yt-dlp dump-single-json
* @param {string} [lang]
* @returns {Array<{ track: any, lang: string|null }>}
*/
export function orderedTracks(json, lang = 'fr') {
const manual = (json && json.subtitles) || {};
const auto = (json && json.automatic_captions) || {};
const seen = new Set();
const out = [];
const push = (key, dict) => {
if (key == null) return;
const t = firstTrack(dict[key]);
if (!t || !t.url || seen.has(String(t.url))) return;
seen.add(String(t.url));
out.push({ track: t, lang: key });
};
const primary = pickTrack(json, lang);
if (primary.track) {
seen.add(String(primary.track.url));
out.push({ track: primary.track, lang: primary.lang });
}
// Original language first fallback (avoids translated-track rate limits)
const reqPrimary = String(lang || '').toLowerCase().split('-')[0];
for (const fallbackLang of ['en', 'fr']) {
if (fallbackLang === reqPrimary) continue;
const k = findLangKey(manual, fallbackLang);
if (k) push(k, manual);
const ka = findLangKey(auto, fallbackLang);
if (ka) push(ka, auto);
}
for (const key of Object.keys(manual)) push(key, manual);
for (const key of Object.keys(auto)) push(key, auto);
return out;
}
/**
* Parse a YouTube `json3` timedtext payload into normalized lines.
* @param {any} data parsed JSON
* @returns {TranscriptLine[]}
*/
export function parseJson3(data) {
const events = (data && data.events) || [];
if (!Array.isArray(events)) return [];
const out = [];
for (const e of events) {
if (!e || !Array.isArray(e.segs)) continue;
const text = e.segs.map((s) => String(s?.utf8 ?? '')).join('').replace(/\n+/g, ' ').trim();
if (!text) continue;
out.push({
t: Number(e.tStartMs || 0) / 1000,
dur: Number(e.dDurationMs || 0) / 1000,
text,
});
}
return capLines(out);
}
/** Decode a handful of HTML entities found in VTT payloads. */
function decodeEntities(s) {
return String(s || '')
.replace(/&amp;/g, '&')
.replace(/&lt;/g, '<')
.replace(/&gt;/g, '>')
.replace(/&quot;/g, '"')
.replace(/&#39;/g, "'")
.replace(/&nbsp;/g, ' ');
}
/** Strip XML/HTML tags and decode entities to plain cue text. */
function xmlTextToPlain(s) {
return decodeEntities(String(s || '').replace(/<[^>]*>/g, ' ').replace(/\s+/g, ' ').trim()).trim();
}
function parseTimeAttrToSeconds(v) {
if (v == null || v === '') return null;
const s = String(v).trim();
if (!s) return null;
if (/^\d+(\.\d+)?$/.test(s)) return Number(s);
const m = /(?:(\d+):)?([0-5]?\d):([0-5]\d)(?:\.(\d{1,3}))?/.exec(s);
if (!m) return null;
return Number(m[1] || 0) * 3600 + Number(m[2]) * 60 + Number(m[3]) + Number((m[4] || '0').padEnd(3, '0')) / 1000;
}
/**
* Parse YouTube XML caption formats (`srv1`/`srv2`/`srv3` `<p>` cues,
* `ttml` `<p begin= end=>` cues) into normalized lines.
* @param {string} text raw XML
* @returns {TranscriptLine[]}
*/
export function parseXmlCaptions(text) {
const out = [];
const src = String(text || '');
if (!src || !/<(p|text|tt|transcript)\b/i.test(src)) return capLines(out);
// YouTube srv*: <p t="0" d="2500">Hello <s>world</s></p>
const srvRe = /<p\b[^>]*?(?:t|start)="([^"]*)"[^>]*?(?:d|dur)="([^"]*)"[^>]*>([\s\S]*?)<\/p\s*>/gi;
// Generic ttml: <p begin="00:00:00.000" end="00:00:02.500">...</p>
const ttmlRe = /<p\b[^>]*?begin="([^"]*)"[^>]*?end="([^"]*)"[^>]*>([\s\S]*?)<\/p\s*>/gi;
let m;
while ((m = srvRe.exec(src)) !== null) {
const start = Number(m[1]) / 1000;
const dur = Number(m[2]) / 1000;
const cleaned = xmlTextToPlain(m[3]);
if (!cleaned || !Number.isFinite(start)) continue;
out.push({ t: start, dur: Number.isFinite(dur) && dur > 0 ? dur : 0, text: cleaned });
}
if (out.length === 0) {
while ((m = ttmlRe.exec(src)) !== null) {
const start = parseTimeAttrToSeconds(m[1]);
const end = parseTimeAttrToSeconds(m[2]);
const cleaned = xmlTextToPlain(m[3]);
if (!cleaned || start == null || end == null) continue;
out.push({ t: start, dur: Math.max(0, end - start), text: cleaned });
}
}
// Plain <text start="2.5" dur="2.0">Fallback</text> variant
if (out.length === 0) {
const textRe = /<text\b[^>]*?start="([^"]*)"[^>]*?(?:dur="([^"]*)")?[^>]*>([\s\S]*?)<\/text\s*>/gi;
while ((m = textRe.exec(src)) !== null) {
const start = Number(m[1]);
const dur = Number(m[2] || 0);
const cleaned = xmlTextToPlain(m[3]);
if (!cleaned || !Number.isFinite(start)) continue;
out.push({ t: start, dur: dur > 0 ? dur : 0, text: cleaned });
}
}
return capLines(out);
}
/** Detect YouTube "Sorry / automated queries" style HTML error pages. */
export function looksLikeHtmlError(text) {
const s = String(text || '').trimStart().slice(0, 512).toLowerCase();
return s.startsWith('<!doctype') || s.startsWith('<html');
}
function vttTimestampToSeconds(h, m, s, ms) {
return Number(h) * 3600 + Number(m) * 60 + Number(s) + Number(ms) / 1000;
}
/**
* Parse a WebVTT payload into normalized lines.
* Handles cue identifiers, multi-line cues and inline tags (`<c>`, `<i>`, timestamps).
* @param {string} text raw VTT
* @returns {TranscriptLine[]}
*/
export function parseVtt(text) {
if (looksLikeHtmlError(text)) return [];
const out = [];
const blocks = String(text || '').replace(/\r\n/g, '\n').split(/\n{2,}/);
const tsRe = /(?:(\d{2,}):)?([0-5]?\d):([0-5]\d)\.(\d{3})\s*-->\s*(?:(\d{2,}):)?([0-5]?\d):([0-5]\d)\.(\d{3})/;
for (const block of blocks) {
const lines = String(block || '').split('\n');
const timingIdx = lines.findIndex((l) => l.includes('-->'));
if (timingIdx === -1) continue;
const m = tsRe.exec(lines[timingIdx]);
if (!m) continue;
const start = vttTimestampToSeconds(m[1] || '0', m[2], m[3], m[4]);
const end = vttTimestampToSeconds(m[5] || '0', m[6], m[7], m[8]);
const content = lines
.slice(timingIdx + 1)
.join(' ')
// strip inline VTT tags (<c.colorE5E5E5>, <i>, <00:00:01.000>, ...)
.replace(/<[^>]*>/g, ' ')
.replace(/\s+/g, ' ')
.trim();
const decoded = decodeEntities(content).trim();
// Skip WEBVTT header leftovers and empty cues
if (!decoded || /^WEBVTT/i.test(decoded)) continue;
out.push({ t: start, dur: Math.max(0, end - start), text: decoded });
}
return capLines(out);
}
/** Bound response volume for very long transcripts (Phase 1 decision: truncate). */
export function capLines(lines, max = MAX_LINES_DEFAULT) {
const list = Array.isArray(lines) ? lines : [];
const n = Math.max(1, Number(max || MAX_LINES_DEFAULT));
return list.length > n ? list.slice(0, n) : list;
}
/**
* Parse a downloaded timedtext payload regardless of its format.
* Tries `json3`, then `vtt`, then XML captions (`srv*`/`ttml`).
* HTML error pages always yield `[]`.
*/
export function parseTrackText(text, ext = '') {
if (looksLikeHtmlError(text)) return [];
const e = String(ext || '').toLowerCase();
const asJson3 = () => {
try { return parseJson3(JSON.parse(String(text))); } catch { return []; }
};
if (e === 'json3') {
const lines = asJson3();
if (lines.length) return lines;
const vtt = parseVtt(text);
if (vtt.length) return vtt;
return parseXmlCaptions(text);
}
const vtt = parseVtt(text);
if (vtt.length) return vtt;
const j = asJson3();
if (j.length) return j;
return parseXmlCaptions(text);
}
/** Map short (`yt`) and long (`youtube`) provider ids to `providerUrlFrom()` names. */
export function normalizeTranscriptProvider(provider) {
const p = String(provider || '').trim().toLowerCase();
const map = {
yt: 'youtube', youtube: 'youtube',
dm: 'dailymotion', dailymotion: 'dailymotion',
tw: 'twitch', twitch: 'twitch',
pt: 'peertube', peertube: 'peertube',
od: 'odysee', odysee: 'odysee',
ru: 'rumble', rumble: 'rumble',
};
return map[p] || null;
}
export { trackExtOf as transcriptTrackExt, firstTrack as transcriptFirstTrack };
+72
View File
@@ -0,0 +1,72 @@
import { Injectable, inject } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { Observable, of } from 'rxjs';
import { tap } from 'rxjs/operators';
import { dedupeSort } from './suggest.util';
export interface SuggestResponse {
q: string;
groups: Record<string, string[]>;
}
interface SuggestCacheEntry {
t: number;
data: SuggestResponse;
}
/**
* Step 15 — Query typeahead service.
*
* Contract: no network while `q.trim().length < 2`. Callers debounce
* (250-300ms) + `switchMap` (in-flight cancellation). Responses are cached
* 5 minutes, deduped and sorted per provider group.
*/
@Injectable({ providedIn: 'root' })
export class SuggestService {
private http = inject(HttpClient);
private cache = new Map<string, SuggestCacheEntry>();
private cacheTtlMs = 5 * 60 * 1000;
private apiBase(): string {
try {
const port = window?.location?.port || '';
return port && port !== '4000' ? '/proxy/api' : '/api';
} catch {
return '/api';
}
}
/** Single suggest round-trip (no debounce here — the caller owns it). */
suggestOnce(q: string, providers?: string[] | 'all', limit = 10): Observable<SuggestResponse> {
const query = String(q ?? '').trim();
if (query.length < 2) return of({ q: query, groups: {} });
const prov = Array.isArray(providers) && providers.length > 0 ? [...providers].sort().join(',') : 'all';
const key = `${prov}|${query.toLowerCase()}|${limit}`;
const now = Date.now();
const hit = this.cache.get(key);
if (hit && now - hit.t < this.cacheTtlMs) return of(hit.data);
const params = new URLSearchParams({ q: query, limit: String(limit) });
if (Array.isArray(providers) && providers.length > 0) params.set('providers', providers.join(','));
return this.http.get<SuggestResponse>(`${this.apiBase()}/search/suggest?${params.toString()}`).pipe(
tap((res) => {
try {
const groups: Record<string, string[]> = {};
for (const [pid, list] of Object.entries(res?.groups || {})) {
groups[pid] = dedupeSort(Array.isArray(list) ? list : [], query).slice(0, limit);
}
const clean: SuggestResponse = { q: res?.q ?? query, groups };
if (this.cache.size > 200) this.cache.clear();
this.cache.set(key, { t: Date.now(), data: clean });
// Mutate in place so subscribers see the cleaned payload
res.q = clean.q;
res.groups = clean.groups;
} catch {}
}),
);
}
clearCache(): void {
this.cache.clear();
}
}
+103
View File
@@ -0,0 +1,103 @@
import 'zone.js/node';
import '@angular/compiler'; // JIT needed: decorators are evaluated at import time under ts-node
import { firstValueFrom, of } from 'rxjs';
import { Injector, runInInjectionContext } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { SuggestService } from './suggest.service';
import { dedupeSort, mergeGroups, highlightParts } from './suggest.util';
/**
* Step 15 — Unit tests for query typeahead (parsing / dedup / min-length guard).
*
* Same harness as Step 11: ts-node isolated transpilation, manual injection
* with a minimal HttpClient fake returning real Observables.
* Run with: npm run test:suggest
*/
function assertEqual(actual: unknown, expected: unknown, message: string): void {
const a = JSON.stringify(actual);
const b = JSON.stringify(expected);
if (a !== b) {
throw new Error(`${message} (expected ${b}, got ${a})`);
}
}
function assert(condition: unknown, message: string): void {
if (!condition) {
throw new Error(message);
}
}
function logOk(msg: string): void {
console.log(`✓ ${msg}`);
}
(async () => {
// --- Pure helpers: dedupe + sort ---
assertEqual(
dedupeSort(['Tutoriel Angular', 'tutoriel angular ', ' ', 'Tutoriel Android'], 'tutoriel an'),
['Tutoriel Android', 'Tutoriel Angular'],
'dedupeSort dedupes case-insensitively and sorts prefix-first then alpha',
);
logOk('dedupeSort dedup + tri');
assertEqual(dedupeSort(['b', 'a', 'c']), ['a', 'b', 'c'], 'dedupeSort sorts alphabetically without query');
logOk('dedupeSort alpha fallback');
// --- Pure helpers: merge groups ---
const merged = mergeGroups({ dm: ['Tutoriel android'], yt: ['Tutoriel Angular', 'tutoriel android'], xx: ['zz'] });
assertEqual(
merged,
[
{ text: 'Tutoriel Angular', provider: 'yt' },
{ text: 'tutoriel android', provider: 'yt' },
{ text: 'zz', provider: 'xx' },
],
'mergeGroups respects registry order, dedupes across providers, unknowns last',
);
logOk('mergeGroups order + dedup');
// --- Pure helpers: highlight ---
assertEqual(
highlightParts('Tutoriel Angular', 'toriel an'),
{ pre: 'Tu', match: 'toriel An', post: 'gular' },
'highlightParts splits on first case-insensitive occurrence',
);
assertEqual(
highlightParts('Bonjour', 'zzz'),
{ pre: 'Bonjour', match: '', post: '' },
'highlightParts without occurrence returns plain text',
);
logOk('highlightParts');
// --- Service: min length guard (no network below 2 chars) ---
let httpCalls = 0;
const httpFake = {
get: (url: string) => {
httpCalls += 1;
assert(url.includes('/search/suggest'), `suggest hits /api/search/suggest (got ${url})`);
return of({ q: 'angular', groups: { yt: ['angular tutorial', 'Angular'], dm: ['angular tutorial'] } });
},
};
const injector = Injector.create([{ provide: HttpClient, useValue: httpFake }]);
const svc: SuggestService = runInInjectionContext(injector, () => new SuggestService());
const short = await firstValueFrom(svc.suggestOnce('a'));
assertEqual(short, { q: 'a', groups: {} }, 'q < 2 chars returns empty without network');
assertEqual(httpCalls, 0, 'no HTTP call for short queries');
logOk('suggestOnce min-length guard (no network)');
// --- Service: dedup + cache ---
const first = await firstValueFrom(svc.suggestOnce('angular', ['yt', 'dm']));
assertEqual(httpCalls, 1, 'one HTTP call for a fresh query');
assertEqual(first.groups['yt'], ['Angular', 'angular tutorial'], 'per-provider groups deduped + sorted');
const second = await firstValueFrom(svc.suggestOnce('angular', ['dm', 'yt']));
assertEqual(httpCalls, 1, 'second identical query served from cache (provider order-insensitive)');
assert(second === first || JSON.stringify(second) === JSON.stringify(first), 'cached payload matches');
logOk('suggestOnce dedup + cache 5 min');
console.log('\nAll suggest tests passed.');
})().catch((e) => {
console.error(e instanceof Error ? e.stack : String(e));
process.exit(1);
});
+72
View File
@@ -0,0 +1,72 @@
/**
* Step 15 — Pure helpers for query typeahead (testable without Angular).
*/
/** Normalize a suggestion string for dedup comparisons. */
export function normSuggest(s: string): string {
return String(s ?? '').trim().toLowerCase().replace(/\s+/g, ' ');
}
/**
* Deduplicate (case-insensitive, first occurrence wins) then sort:
* items starting with `q` come first, alphabetical within each group.
*/
export function dedupeSort(items: string[], q = ''): string[] {
const seen = new Set<string>();
const clean: string[] = [];
for (const raw of items || []) {
const t = String(raw ?? '').trim();
if (!t) continue;
const k = normSuggest(t);
if (seen.has(k)) continue;
seen.add(k);
clean.push(t);
}
const nq = normSuggest(q);
clean.sort((a, b) => {
const na = normSuggest(a);
const nb = normSuggest(b);
const pa = nq ? (na.startsWith(nq) ? 0 : 1) : 0;
const pb = nq ? (nb.startsWith(nq) ? 0 : 1) : 0;
if (pa !== pb) return pa - pb;
return na < nb ? -1 : na > nb ? 1 : 0;
});
return clean;
}
/**
* Flatten provider groups into a single ordered list (registry order),
* deduped case-insensitively. Unknown provider keys go last (insertion order).
*/
export function mergeGroups(
groups: Record<string, string[]> | null | undefined,
order: string[] = ['yt', 'dm', 'tw', 'pt', 'od', 'ru'],
): Array<{ text: string; provider: string }> {
const out: Array<{ text: string; provider: string }> = [];
const seen = new Set<string>();
const push = (pid: string, list: string[] | undefined) => {
for (const raw of list || []) {
const t = String(raw ?? '').trim();
if (!t) continue;
const k = normSuggest(t);
if (seen.has(k)) continue;
seen.add(k);
out.push({ text: t, provider: pid });
}
};
for (const pid of order) push(pid, groups?.[pid]);
for (const pid of Object.keys(groups || {})) {
if (!order.includes(pid)) push(pid, (groups as Record<string, string[]>)[pid]);
}
return out;
}
/** Split `text` on the first case-insensitive occurrence of `q` for <mark> highlighting. */
export function highlightParts(text: string, q: string): { pre: string; match: string; post: string } {
const t = String(text ?? '');
const needle = String(q ?? '').trim();
if (!needle) return { pre: t, match: '', post: '' };
const idx = t.toLowerCase().indexOf(needle.toLowerCase());
if (idx === -1) return { pre: t, match: '', post: '' };
return { pre: t.slice(0, idx), match: t.slice(idx, idx + needle.length), post: t.slice(idx + needle.length) };
}
+269 -7
View File
@@ -23,6 +23,7 @@ import { formatAbsoluteFr, formatNumberFr } from '../../utils/date.util';
import { ProviderBadgeComponent } from '../../app/shared/components/provider-badge/provider-badge.component';
import { ChannelIdentityComponent } from '../../app/shared/components/channel-identity/channel-identity.component';
import { DurationPipe } from '../../app/shared/pipes/duration.pipe';
import { UserService } from '../../services/user.service';
@Component({
selector: 'app-watch',
@@ -44,6 +45,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
private auth = inject(AuthService);
private http = inject(HttpClient);
private iframeProgress = inject(IframeProgressService);
private users = inject(UserService);
private channels = inject(ChannelsService);
private subs = inject(SubscriptionsService);
private routeSubscription!: Subscription;
@@ -197,6 +199,8 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
// Build a direct Twitch URL as a fallback open-in-new-tab action
twitchOpenUrl(): string | null {
if (this.provider() !== 'twitch') return null;
const clip = this.twitchClip();
if (clip) return `https://clips.twitch.tv/${encodeURIComponent(clip)}`;
const ch = this.twitchChannel();
const id = this.videoId();
const cur = this.video();
@@ -286,6 +290,20 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
liked = signal<boolean>(false);
likeBusy = signal<boolean>(false);
// --- Transcript state (Step 16, Phase 1: display + language selector, no seek) ---
transcriptOpen = signal(false);
transcriptLoading = signal(false);
transcriptAvailable = signal(false);
transcriptError = signal<string | null>(null);
// Machine-readable unavailability reason from the API: 'no_subtitles' |
// 'provider_unsupported' | 'temporarily_unavailable' | null.
transcriptReason = signal<string | null>(null);
// True when retrying may succeed (transient YouTube rate-limit/fetch failure).
transcriptRetryable = signal(false);
transcriptLanguages = signal<string[]>([]);
transcriptLines = signal<Array<{ t: number; dur: number; text: string }>>([]);
selectedTranscriptLang = signal('fr');
// --- Watch history tracking for native player ---
watchHistoryId = signal<string | null>(null);
startPositionSeconds = signal<number | null>(null);
@@ -296,6 +314,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
private providerSel = signal<string>('');
provider = computed(() => this.providerSel() || this.instances.selectedProvider());
private twitchChannel = signal<string | null>(null);
private twitchClip = signal<string | null>(null);
// Backend-provided embed URL for providers that need it (e.g., Rumble)
private rumbleEmbed = signal<string | null>(null);
@@ -303,9 +322,10 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
const id = this.videoId();
const p = this.provider();
const ch = this.twitchChannel();
const clip = this.twitchClip();
const slug = this.odyseeSlug();
const start = Math.max(0, this.startPositionSeconds() || 0);
if (!id && !(p === 'twitch' && ch) && !(p === 'odysee' && slug)) return null;
if (!id && !(p === 'twitch' && (ch || clip)) && !(p === 'odysee' && slug)) return null;
try {
const host = (location && location.hostname) ? location.hostname : 'localhost';
if (p === 'youtube') {
@@ -324,9 +344,15 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
return this.sanitizer.bypassSecurityTrustResourceUrl(u);
}
if (p === 'twitch') {
// Channel or VOD; Twitch requires one or more 'parent' params that match the embedding domain
// Channel, clip ou VOD ; Twitch exige 'parent' = domaine d'embed.
const parentsSet = new Set<string>([host, 'localhost', '127.0.0.1']);
const parentParams = Array.from(parentsSet).map(h => `parent=${encodeURIComponent(h)}`).join('&');
// Clips : player dédié clips.twitch.tv (le player VOD rejette les slugs).
const clipId = clip || (/^\d+$/.test(String(id || '')) ? null : ((this.video() as any)?.kind === 'clip' ? id : null));
if (clipId) {
const u = `https://clips.twitch.tv/embed?clip=${encodeURIComponent(clipId)}&${parentParams}&autoplay=false`;
return this.sanitizer.bypassSecurityTrustResourceUrl(u);
}
const base = 'https://player.twitch.tv/';
// ?video= exige un id de VOD numérique ; un login (non numérique) sans
// ?channel= doit être traité comme une chaîne (résultats de recherche).
@@ -461,11 +487,13 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
const slug = this.route.snapshot.queryParamMap.get('slug');
const p = (this.route.snapshot.queryParamMap.get('p') || providerFromPath) as Provider | null;
const ch = this.route.snapshot.queryParamMap.get('channel');
const clip = this.route.snapshot.queryParamMap.get('clip');
if (id) {
this.videoId.set(id);
this.odyseeSlug.set(slug);
this.providerSel.set(p || '');
this.twitchChannel.set(ch);
this.twitchClip.set(clip);
// Choisir un provider final (paramètre présent, segment d'URL, ou fallback depuis état courant)
let finalProvider = p as Provider | null;
@@ -583,6 +611,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
this.summaryError.set(null);
this.selectedQuality.set(null);
this.resetDownloadUi();
this.resetTranscriptUi();
// Reset like state while loading
this.liked.set(false);
this.likeBusy.set(false);
@@ -861,6 +890,189 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
this.selectedQuality.set(value || null);
}
// --- Transcript logic (Step 16) ---
private resetTranscriptUi() {
this.transcriptOpen.set(false);
this.transcriptLoading.set(false);
this.transcriptAvailable.set(false);
this.transcriptError.set(null);
this.transcriptReason.set(null);
this.transcriptRetryable.set(false);
this.transcriptLanguages.set([]);
this.transcriptLines.set([]);
try {
const prefLang = String(this.users.preferences()?.language || '').trim().toLowerCase().slice(0, 12);
this.selectedTranscriptLang.set(prefLang || 'fr');
} catch {
this.selectedTranscriptLang.set('fr');
}
}
toggleTranscript() {
const next = !this.transcriptOpen();
this.transcriptOpen.set(next);
if (next && this.transcriptLines().length === 0 && !this.transcriptLoading()) {
this.loadTranscript();
}
}
onTranscriptLangChange(value: string) {
const lang = String(value || '').trim().slice(0, 12) || 'fr';
this.selectedTranscriptLang.set(lang);
this.loadTranscript();
}
/** Map a backend transcript state (reason/error/retryable) to the UI:
* user-facing message + whether a "Réessayer" button makes sense. */
private transcriptStatusFor(data: any): { message: string | null; retryable: boolean; reason: string | null } {
const reason = typeof data?.reason === 'string' ? data.reason : null;
const code = typeof data?.error === 'string' ? data.error : '';
const retryable = data?.retryable === true
|| code === 'transcript_temporarily_unavailable'
|| code === 'rate_limited';
if (reason === 'provider_unsupported') {
return {
message: 'Les transcripts ne sont pas pris en charge pour ce fournisseur (aucun sous-titre exposé).',
retryable: false,
reason,
};
}
if (reason === 'temporarily_unavailable' || code === 'transcript_temporarily_unavailable') {
return {
message: 'Sous-titres temporairement indisponibles (limite YouTube). Réessayez dans quelques instants.',
retryable: true,
reason: reason || 'temporarily_unavailable',
};
}
if (reason === 'no_subtitles') {
return { message: null, retryable: false, reason };
}
// Legacy / unknown error code: sanitize, never leak HTML or parse noise.
const clean = this.transcriptDisplayError(code);
if (clean) return { message: `Transcript indisponible. (${clean})`, retryable, reason };
if (retryable) {
return {
message: 'Sous-titres temporairement indisponibles. Réessayez dans quelques instants.',
retryable: true,
reason: reason || 'temporarily_unavailable',
};
}
return { message: null, retryable: false, reason };
}
/** Sanitize a backend error code for display: never leak HTML pages or raw
* JSON-parse noise (e.g. proxy / CDN "<!DOCTYPE..." bodies surfaced as
* "SyntaxError: Unexpected token '<'..."). Returns null when nothing
* user-friendly can be shown (the "no subtitles" notice already covers it). */
private transcriptDisplayError(raw: unknown): string | null {
let msg = '';
try {
if (typeof raw === 'string') msg = raw;
else if (raw instanceof Error) msg = raw.message || '';
else if (raw != null) msg = String(raw);
} catch {
return null;
}
msg = msg.trim().slice(0, 160);
if (!msg) return null;
const lowered = msg.toLowerCase();
// Ignore internal codes handled silently + any HTML / JSON-parse noise.
if (
msg === 'transcript_fetch_failed' ||
msg === 'transcript_temporarily_unavailable' ||
msg === 'rate_limited' ||
lowered.includes('<!doctype') ||
lowered.includes('<html') ||
msg.includes('<') ||
lowered.includes('not valid json') ||
lowered.includes('unexpected token') ||
lowered.startsWith('syntaxerror')
) {
return null;
}
return msg;
}
retryTranscript() {
if (this.transcriptLoading()) return;
this.transcriptLines.set([]);
this.loadTranscript();
}
loadTranscript() {
const p = this.provider();
const id = this.videoId();
if (!p || !id) {
this.transcriptError.set('Vidéo introuvable.');
this.transcriptReason.set(null);
this.transcriptRetryable.set(false);
return;
}
// Never break the Watch page: every failure degrades to the "unavailable" notice.
try {
this.transcriptLoading.set(true);
this.transcriptError.set(null);
this.transcriptReason.set(null);
this.transcriptRetryable.set(false);
const opts = this.buildProviderOpts() as Record<string, string>;
const params = new URLSearchParams({ lang: this.selectedTranscriptLang() });
if (opts['instance']) params.set('instance', opts['instance']);
if (opts['slug']) params.set('slug', opts['slug']);
const url = `${this.apiBase()}/transcript/${encodeURIComponent(p)}/${encodeURIComponent(id)}?${params.toString()}`;
this.http.get<any>(url).subscribe({
next: (data) => {
this.transcriptLoading.set(false);
this.transcriptAvailable.set(!!data?.available);
this.transcriptLanguages.set(Array.isArray(data?.languages) ? data.languages : []);
this.transcriptLines.set(Array.isArray(data?.lines) ? data.lines : []);
if (data?.available && typeof data?.lang === 'string' && data.lang) {
this.selectedTranscriptLang.set(data.lang);
}
if (!data?.available) {
const st = this.transcriptStatusFor(data);
this.transcriptReason.set(st.reason);
this.transcriptRetryable.set(st.retryable);
this.transcriptError.set(st.message);
} else {
this.transcriptReason.set(null);
this.transcriptRetryable.set(false);
this.transcriptError.set(null);
}
},
error: (err) => {
this.transcriptLoading.set(false);
this.transcriptAvailable.set(false);
this.transcriptLines.set([]);
const body = err?.error;
if (body && typeof body === 'object' && Array.isArray(body.languages)) {
this.transcriptLanguages.set(body.languages);
}
const raw =
(body && typeof body === 'object' && (body.error || body.details || body.reason)) ||
(typeof body === 'string' ? body : '') ||
err?.message ||
'';
const status = err?.status;
if (status === 429 || String(raw || '').trim() === 'rate_limited') {
this.transcriptReason.set('temporarily_unavailable');
this.transcriptRetryable.set(true);
this.transcriptError.set('Trop de requêtes. Réessayez dans une minute.');
return;
}
const st = this.transcriptStatusFor(
body && typeof body === 'object' ? body : { error: raw },
);
this.transcriptReason.set(st.reason);
this.transcriptRetryable.set(st.retryable);
this.transcriptError.set(st.message);
},
});
} catch {
this.transcriptLoading.set(false);
this.transcriptAvailable.set(false);
}
}
// --- Download logic ---
openDownloadPanel() {
this.downloadOpen.set(true);
@@ -1053,6 +1265,10 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
if (slug.startsWith('/')) slug = slug.slice(1);
qp.slug = slug;
}
if (p === 'twitch' && String((v as any).kind || '').toLowerCase() === 'clip') {
qp.clip = v.videoId;
return qp;
}
if (p === 'twitch' && (v.type === 'channel' || (v as any).type === 'live')) {
// Prefer explicit videoId as channel login; fallback to parsing URL
if (v.videoId && !/^\d+$/.test(String(v.videoId))) {
@@ -1072,18 +1288,64 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
private loadRelatedSuggestions(): void {
const cur = this.video();
if (!cur) return;
const title = (cur.title || '').trim();
const maxItems = 12;
const apply = (items: any[]) => {
// Ignore les réponses arrivées après une navigation (race).
try {
if ((this.videoId() || '') !== (cur.videoId || '')) return;
} catch {}
const list = (items || []).filter((v: any) => (v?.videoId || v?.id) !== cur.videoId).slice(0, maxItems);
this.video.update(v => v ? { ...v, relatedStreams: list } : v);
};
const titleFallback = () => {
const title = (cur.title || '').trim();
if (title) {
this.apiService.searchVideosPage(title).subscribe(res => {
const items = (res.items || []).filter(v => v.videoId !== cur.videoId).slice(0, maxItems);
this.video.update(v => v ? { ...v, relatedStreams: items } : v);
apply(res.items || []);
});
} else {
this.apiService.getTrendingPage().subscribe(res => {
const items = (res.items || []).filter(v => v.videoId !== cur.videoId).slice(0, maxItems);
this.video.update(v => v ? { ...v, relatedStreams: items } : v);
apply(res.items || []);
});
}
};
// YouTube : vrai watch-next InnerTube via /api/details (0 quota, 0 recherche
// supplémentaire) — façon SmartTube. Autres providers : fallback existant.
try {
if (this.provider() === 'youtube' && cur.videoId) {
const id = cur.videoId;
this.http.get<any>(`/api/details/youtube/${encodeURIComponent(id)}`).subscribe({
next: (data: any) => {
const rel = Array.isArray(data?.related) ? data.related : [];
if (rel.length > 0) {
apply(rel.map((it: any) => ({
url: it.url || `https://www.youtube.com/watch?v=${it.id}`,
type: 'video',
title: it.title || '',
thumbnail: it.thumbnail || '',
uploaderName: it.uploaderName || '',
uploaderAvatar: '',
channelId: it.channelId || undefined,
channelHandle: it.channelHandle || undefined,
channelExternalId: it.channelExternalId || it.channelId || undefined,
uploadedDate: it.publishedAt || '',
duration: typeof it.duration === 'number' ? it.duration : 0,
views: typeof it.views === 'number' ? it.views : 0,
uploaded: 0,
videoId: String(it.id || ''),
provider: 'youtube',
isShort: it.isShort === true ? true : undefined,
isLive: it.isLive === true ? true : undefined,
} as Video)));
} else {
titleFallback();
}
},
error: () => titleFallback(),
});
return;
}
} catch {}
titleFallback();
}
}
+282
View File
@@ -17,6 +17,179 @@
- [x] Step 12: Write basic e2e scenarios for providers filtering and deep-link
- [x] Step 13: Add minimal telemetry hooks and events
- [x] Step 14: Update README with search UX section, keyboard shortcuts and usage notes
- [x] Step 15: Auto-complétion de la **requête** (typeahead) dans la barre de recherche
- [x] Step 16: **Transcripts** de vidéos sur la page Watch (API + UI, Phase 1 du doc)
---
## Step 15 — Suggestions de requêtes dans la barre de recherche
### Ce que ça veut dire (et ce que ce n'est pas)
Quand l'utilisateur tape `tutoriel an`, la barre affiche **avant qu'il valide** une liste de
requêtes complètes probables : `tutoriel angular`, `tutoriel android`… + ses recherches passées.
> ⚠️ À ne pas confondre avec `SearchSuggestionsComponent` (`src/components/search/search-suggestions.component.ts`),
> qui lui affiche les **résultats de recherche** groupés par provider **après** la validation sur `/search`.
> Le Step 15 ajoute un nouveau panneau de **suggestions de texte** sous l'`<input>`.
### Ce qui existe déjà (à réutiliser)
| Élément | Fichier | État |
|---|---|---|
| `@Output() searchChange` (query debouncée 300 ms) | `src/components/search/search-box.component.ts:46` | Émis mais **aucun abonné** — le header ne branche que `(submitted)` (`header.component.html:19`) |
| Popover `@provider` + navigation clavier | `search-box.component.ts:16` (`AT_QUERY_RE`), `.html:56-76` | ✅ mais ne couvre que les ids de providers |
| Historique de recherches | `HistoryService.getSearchHistory(n)` | ✅ utilisé dans le Quick Menu |
| Cache in-memory TTL 60 s | `src/app/search/search.service.ts:33` | Pattern à dupliquer |
| Fan-out multi-providers | `GET /api/search` (`server/index.mjs:2116`) | Pattern à dupliquer |
### Livrables
1. **Backend** — `GET /api/search/suggest?q=…&providers=yt,dm&limit=10`
→ `{ q, groups: { yt: string[], dm: string[], … } }`
- Un `suggest(q, { limit })` optionnel par handler dans `server/providers/*.mjs` + déclaré dans
`server/providers/registry.mjs` (`ProviderAdapter`).
- Providers sans API de suggestions native → `[]` (dégradation propre, jamais d'erreur 500).
- Cache + rate-limit sur le même modèle que `/api/search`.
2. **Front service** — `SuggestService` (ou méthode `suggest()` sur `SearchService`) :
`min length 2`, `debounce 250-300 ms`, `switchMap` (annulation de la requête précédente),
cache 5 min, dédoublonnage + tri.
3. **UI** — panneau sous l'input, sections :
- 🔎 « Recherches récentes » (History, avec icône horloge)
- 📊 Suggestions providers, groupées par provider (badges `YT/DM/…`) ou fusionnées
- Sous-chaîne commune mise en évidence
4. **Interactions** — clic = remplir + soumettre ; ↑/↓/Enter/Tab/Esc ; **priorité au popover `@`**
quand il est ouvert ; `aria-listbox` + `aria-activedescendant` (le `role="combobox"` est déjà sur l'input).
5. **Télémétrie** — `suggest_shown`, `suggest_used` (à ajouter à la whitelist serveur de l'étape 13).
6. **Tests** — `npm run test:suggest` (unitaires parsing/dédup) + scénario e2e : `providers=yt` ne renvoie
que des suggestions `yt`, provider inconnu → fallback registry complet.
### Critères d'acceptation
- [x] Aucune requête réseau tant que `q.trim().length < 2`
- [x] frappe rapide = une seule requête en vol (pas d'empilement / de réponse qui écrase la dernière)
- [x] Esc ferme le panneau sans effacer le texte ; Tab complète sans lancer la recherche
- [x] Un provider sans suggestions ou en échec n'affiche ni trou ni erreur dans le panneau
- [x] Focus clavier et lecteurs d'écran cohérents avec l'existant (a11y Step 10)
---
## Step 16 — Transcripts (Phase 1 du `docs/transcript_architecture_dev.md`)
### Principe directeur
> **Un endpoint, un module de parsing, une UI.** La disponibilité dépend de la plateforme, mais le
> **code ne change pas** selon le provider (le doc appelle ça « Rung 2 : réutiliser, pas réinventer »).
Source des données : `yt-dlp --dump-single-json --skip-download` (déjà utilisé côté serveur, ex.
`server/index.mjs:1012`) → on lit `subtitles` / `automatic_captions` → on télécharge la piste `json3` ou `vtt`
→ on normalise en `{ lang, available, languages, lines: [{ t, dur, text }] }`.
### Périmètre Phase 1 (à faire)
**Backend**
- [x] `server/transcript.mjs` : fonctions **pures** `pickTrack(json, lang)`, `parseJson3(data)`, `parseVtt(text)`
(testables sans réseau ni DB)
- [x] Route `GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=` dans `server/index.mjs`
(`instance/slug/sourceUrl` nécessaires pour PeerTube multi-instances ; réutiliser `providerUrlFrom()`).
- [x] `transcriptCache` : `Map` + TTL long (24 h), clé `transcript:${provider}:${videoId}:${lang}`.
- [x] Rate-limit (~10 req/min/IP) sur le motif `channelsLimiter`.
- [x] Aucun sous-titre → **HTTP 200** `{ available: false }` (pas une erreur).
- [x] Gestion 429 persistant + fallback « temp dir » (`writeAutoSub`) et variables d'env
(`TRANSCRIPT_CACHE_TTL`, `TRANSCRIPT_RATE_LIMIT`, `TRANSCRIPT_FALLBACK_TMP`).
**Frontend (variante UI #1 du doc)**
- [x] Bouton « Transcript » à côté du bouton Download (`watch.component.html:159`).
- [x] Panneau copié sur le motif du Download Panel (`watch.component.html:185`) :
`toggleTranscript()` + `loadTranscript()`, chargement/absent/lines, `<select>` de langue
(langue par défaut = préférence utilisateur `language`).
- [x] Le panneau ne doit **jamais** casser la page Watch (providers sans sous-titres : Twitch, Odysee, Rumble).
**Tests & docs**
- [x] `server/tests/transcript.test.mjs` : fixtures `json3`/`vtt` offline (parseurs) + intégrité du contrat API.
- [x] Script `npm run test:transcript` + l'ajouter à `.github/workflows/ci.yml`.
- [x] README : cocher `⏳ Sous-titres & transcripts` dans la roadmap.
### Hors périmètre Phase 1 (à explicitement garder pour plus tard)
- **Clic sur une ligne = seek** (variante UI #2) : exige de brancher par provider
`VideoPlayerComponent.seekBy()` (`src/components/video-player/video-player.component.ts:112`) et
`IframeProgressService.seekTo()` (`src/services/iframe-progress.service.ts:104`, iframe YouTube).
- Cache Redis / multi-instance, proxies rotatifs, monitoring des 429, formats exotiques.
### Décisions à trancher avant de coder
1. **Providers réellement supportés** en Phase 1 (le doc teste YouTube + Dailymotion ; PeerTube a sa propre API
de sous-titres — passer par `yt-dlp` ou par l'API instance ?).
2. **Persistance**: cache mémoire seulement, ou table SQLite `transcripts` (comme pour les téléchargements)
pour survivre aux redémarrages ?
3. **Transcripts générés (auto-captions)** : inclus d'emblée ou option utilisateur ?
4. **Volume de réponse** : les gros transcripts (2 h+) nécessitent-ils une troncature / pagination ?
### Critères d'acceptation
- [x] 1 appel API, 1 parsing, 1 UI identiques quel que soit le provider
- [x] 2ᵉ requête sur la même vidéo = réponse depuis le cache (pas d'appel `yt-dlp`)
- [x] `yt-dlp` en échec → 502 `{ available: false, error: 'transcript_fetch_failed' }`, Watch intacte
- [x] `npm run test:transcript` passe hors ligne (fixtures)
### Décisions Phase 1 (tranchées à l'implémentation)
1. **Providers** : générique via `yt-dlp` (ids courts `yt/dm/…` + longs normalisés), aucun code par plateforme.
2. **Persistance** : cache mémoire seul (TTL 24 h, 200 entrées LRU) ; SQLite reporté en Phase 3.
3. **Auto-captions** : incluses d'emblée (fallback après les sous-titres manuels).
4. **Volume** : troncature `TRANSCRIPT_MAX_LINES` (défaut 5000 lignes).
---
## Idées de nouvelles fonctions utiles
Classées par effort (~ = demi-journée, + = 1-2 jours, ++ = une semaine+). Triées par rapport
valeur / coût pour l'existant.
### Recherche & découverte
| # | Fonction | Description | Effort |
|---|---|---|---|
| 1 | **Filtres de recherche** | Durée (< 4 min / 4-20 / 20+), type (vidéo/chaîne/playlist), date de mise en ligne, langue. `SearchService.sort$` existe déjà, il manque les params + l'UI en chips + le mapping par adapter | + |
| 2 | **Pagination / infinite scroll** | ✅ DONE (Step 18) : backend pages illimitées via continuations InnerTube (`page=2,3…`, `pageSize` ≤ 50) + front `loadNextPage()` + `<app-infinite-anchor>` + `mergeGroups` + `endReached` (`search.component.ts`) | ~ |
| 3 | **Dédup multi-provider** | La même vidéo existe souvent sur Rumble/Odysee/PeerTube : regrouper par titre+durée+chaîne et afficher « aussi disponible sur… » | + |
| 4 | **Recherche dans les résultats** (« search within results ») | Relancer la requête en restreignant au provider/à la chaîne déjà affichée | ~ |
| 5 | **Écran « Recherches récentes »** | Gérer (renommer/supprimer/vider) l'historique de recherche aujourd'hui visible seulement dans le Quick Menu | ~ |
### Watch & lecture
| # | Fonction | Description | Effort |
|---|---|---|---|
| 6 | **Recherche dans le transcript** (`Ctrl+F` du panneau) | Surligne les occurrences + liste des lignes correspondantes ; prolongement naturel du Step 16 | ~ |
| 7 | **Export du transcript** (`.vtt` / `.srt` / `.txt`) | Les lignes `{t, dur, text}` sont déjà normalisées, il ne manque qu'un rendu côté serveur | ~ |
| 8 | **Résumé IA du transcript** | `GEMINI_API_KEY` + `/api/ai/summarize` (commentaires) existent déjà (`server/index.mjs:2050`) : même pattern appliqué au transcript | + |
| 9 | **Chapitres** | Parser `chapters` de `yt-dlp --dump-single-json` (ou les timestamps du titre/description) et les afficher sous le player | + |
| 10 | **Deep-link avec timestamp** | `#/watch/yt/VIDEO?t=123` : seek au démarrage + bouton « Partager à ce moment » | ~ |
| 11 | **Reprendre la lecture** | Mémoriser la position par vidéo dans `history` et proposer « Reprendre à 12:34 » | ~ |
| 12 | **Sélecteur de qualité + « auto » intelligent** | Déjà à la roadmap ; `formatListFromMeta()` et le `format` id existent dans la file de téléchargement, réutilisables pour le player | + |
### Contenus & social
| # | Fonction | Description | Effort |
|---|---|---|---|
| 13 | **Abonnements (chaînes)** | À la roadmap. Le socle est prêt : `server/providers/channel-registry.mjs` + `channel-content.mjs` (videos/shorts/playlists/live). Il manque table `subscriptions`, CRUD API et page « flux d'abonnements » | ++ |
| 14 | **Tags & recherche par tags** | Extraire les tags depuis `dumpSingleJson` et permettre `#/search?tags=angular` | + |
| 15 | **Import/Export playlists** (JSON / OPML) | À la roadmap ; simple contrat de sérialisation + validateur serveur | ~ |
| 16 | **Notifications « nouvelle vidéo »** | Pour les abonnements : poll planifié côté serveur + badge dans le header | + |
### Ops & qualité
| # | Fonction | Description | Effort |
|---|---|---|---|
| 17 | **`/healthz` + page Admin** | À la roadmap : état des clés API (OK/KO), version de `yt-dlp`, hit/miss des caches, compteurs de rate-limit | + |
| 18 | **Cache serveur configurable par provider** | Remplacer le `YT_CACHE_TTL_MS` global par un TTL par provider (`CACHE_TTL_MS_YT=…`) et exposer les clés de cache vides/pleines | ~ |
| 19 | **PWA installable** | Manifest + Service Worker cache des métadonnées (les vidéos restent online) | + |
| 20 | **Superposition raccourcis clavier (`?`)** | Un `@HostListener('document:keydown')` global + modale récapitulative ; tous les raccourcis existent déjà, rien n'est découvrable | ~ |
| 21 | **i18n élargi** | Le pipeline de traduction existe (utilisé comme `'search.placeholder'` + pipeline `t`) ; il reste à compléter les lexiques et sortir les chaînes en dur du FR/EN mélangé actuel | ++ |
## Test commands
```bash
@@ -24,6 +197,8 @@ npm run test:search # Step 11 — unit tests: SearchService + @ parsing +
npm run test:search-e2e # Step 12 — e2e scenarios against a real isolated server
npm run test:preferences # Step 9 — defaultProviders persistence
npm run test:telemetry # Step 13 — telemetry events (insert/list/count)
npm run test:suggest # Step 15 — typeahead parsing/dedup + /api/search/suggest contract
npm run test:transcript # Step 16 — json3/vtt parsers + /api/transcript contract
```
## Implementation notes
@@ -44,3 +219,110 @@ npm run test:telemetry # Step 13 — telemetry events (insert/list/count)
- **Step 14**: README gained a "Recherche unifiée" section (chips, @autocomplete, Ctrl/⌘+K picker, deep-links,
preference fallback, a11y), API endpoints, test commands and roadmap updates. GIF placeholder TODO added.
- CI (`.github/workflows/ci.yml`) now runs preferences, telemetry, search unit and e2e tests.
---
## Step 17 — Plan Anti-Quota YouTube (sans yattee-server en dépendance critique)
### Contexte / diagnostic
- Aujourd'hui `server/providers/youtube.mjs:29-93` + `server/index.mjs:880-1069` utilisent **uniquement la YouTube Data API v3** (`search.list` + `videos.list`) avec rotation `YOUTUBE_API_KEYS` / `YOUTUBE_API_KEY` sur `quotaExceeded|rateLimitExceeded|dailyLimitExceeded|API_KEY_INVALID`.
- Coût officiel : `search.list` = **100 unités**, `videos.list` = **1 unité**, quota gratuit = **10 000 unités/jour/projet** → ~100 recherches/jour max. Chaque page `pageToken` et chaque enrichissement `videos.list` aggrave.
- Cache actuel : `YT_CACHE_TTL_MS` défaut 5 min (`server/index.mjs:904-915`) en mémoire seule, pas de persistance, pas de monitoring de conso.
- Bonne nouvelle : NewTube utilise déjà `yt-dlp` (`youtube-dl-exec`, `YT_DLP_PATH`, `server/index.mjs:83-124`) pour `transcript` + `downloads` + enrichissement Watch. Le `suggest` YT est déjà **sans clé** (`suggestqueries.google.com`, `youtube.mjs:248-264`).
- Conclusion analyse `yattee/yattee-server` : bonne architecture à copier (couches InnerTube → Invidious → yt-dlp + cache + egress-proxy), mais **ne pas l'ajouter comme service critique** (projet de 02/2026, ~107 stars, même combat anti-ban IP/cookies/PO-Token, redondant avec ton yt-dlp direct). Option sidecar seulement en Phase 4.
### Principe cible
> **Primaire sans clé (yt-dlp / InnerTube), API officielle en fallback payant, cache SQLite + mémoire, anti-ban configurable, observabilité.**
```
Front /api/search?providers=yt → searchCache (mémoire + SQLite)
→ YT_SEARCH_MODE=scrape-first (défaut) : yt-dlp `ytsearchN:` / InnerTube → OK ? return
→ sinon fallback API officielle (rotation clés) → OK ? return
→ sinon 200 avec `errors.yt` (dégradation propre, jamais 500)
```
### Phase 0 — Quick wins (0,5 j, sans changer de source)
- [ ] Monter `YT_CACHE_TTL_MS` à 30-60 min pour `yt`, ajouter `SUGGEST_CACHE_TTL_MS` déjà à 5 min, clé cache `q|limit|page|sort`.
- [ ] Réduire le coût API : `maxResults` max 25 au lieu de 50, ne faire `videos.list` que si `details=true`, debounce front déjà 250-300 ms + `switchMap`.
- [ ] Ajouter `GET /healthz` + compteur quota estimé (`search*100 + videos*1`) exposé pour Admin (prépare idée #17).
- [ ] Doc `.env.example` : expliquer `YOUTUBE_API_KEYS` CSV vs JSON, rotation auto.
### Phase 1 — Recherche YT sans clé (2-3 j, cœur du plan)
- [ ] Nouveau `server/providers/youtube-scrape.mjs` :
- `search(q, {limit, page, sort})` via `yt-dlp --dump-single-json --flat-playlist "ytsearch{limit}:{q}"` (timeout 15-20 s, `YT_DLP_PATH` réutilisé), mapping vers `Suggestion` identique à `youtube.mjs:199-232` (title/id/thumbnail/duration/views/publishedAt/channelId/embeddable).
- Support `sort` : `relevance|date|views` → préfixe `ytsearch` + tri local si besoin.
- Jamais de throw bloquant : erreur → throw avec `ytStatus` pour que `/api/search` mette `errors.yt`.
- [ ] Modifier `server/providers/youtube.mjs` en dispatcher :
- `env YT_SEARCH_MODE=scrape-first|api-first|scrape-only|api-only` (défaut `scrape-first`).
- `scrape-first` : essaie scrape, fallback API. `api-first` : comportement actuel. Permet rollback instantané.
- [ ] Cache persistant : table SQLite `youtube_search_cache(q_hash, payload, created_at)` TTL `YT_SCRAPE_TTL_MS` (défaut 30 min) + garde mémoire actuelle. Survit au restart, tue 80% des appels doublons.
- [ ] Rate-limit + concurrence : réutiliser `express-rate-limit` existant, timeout fan-out par provider 8-12 s, `Promise.allSettled` déjà en place dans `/api/search`.
- [ ] Tests : `server/tests/youtube-scrape.test.mjs` offline (fixtures `yt-dlp --flat-playlist` mockées) + e2e `YT_SEARCH_MODE=scrape-only` sans clé → `groups.yt` non vide ; `api-only` sans clé → `errors.yt=youtube_api_key_unavailable`.
- [ ] Env : `YT_SEARCH_MODE`, `YT_SCRAPE_TTL_MS`, `YT_DLP_TIMEOUT_MS`, `YT_DLP_PATH` (déjà), `YOUTUBE_API_KEYS` devient optionnel.
### Phase 2 — Robustesse anti-ban (1-2 j, inspiré yattee-server)
- [ ] Support `YT_COOKIES_FILE` + `YT_PO_TOKEN` passés à yt-dlp (`--cookies`, `--extractor-args youtube:po_token=...`). Doc : comment exporter cookies fresh, rotation manuelle. Sans ça : `Sign in to confirm you're not a bot`.
- [ ] Support `YT_EGRESS_PROXY` (HTTP/SOCKS) pour tout le trafic YT (yt-dlp + fetch InnerTube), configurable au runtime comme yattee `SSRF_EXTRA_ALLOWED_CIDRS` si Invidious LAN.
- [ ] Auto-update yt-dlp : script `npm run ytdlp:update` + check version dans `/healthz` (idée #17). `deno`/`ffmpeg-static` déjà requis pour challenge JS.
- [ ] Observabilité Admin (idée #17) : page `Admin > YouTube` : mode actif, version yt-dlp, hit/miss cache, erreurs `quotaExceeded` vs `bot-check` vs `timeout`, état chaque clé `...abcd OK/KO`.
### Phase 3 — Fonctionnalités bonus débloquées par le sans-clé (1 j / feature)
- [ ] `Trending YT sans clé` : `yt-dlp --flat-playlist "https://www.youtube.com/feed/trending"` → alimente Accueil `Tendances & Viral` sans quota.
- [ ] `Channel browsing enrichi` : réutiliser `channel-registry.mjs` + `channel-content.mjs` via scrape (`videos/shorts/streams/playlists`) au lieu de `channels.list` payant → prérequis Abonnements (idée #13).
- [ ] `Chapitres` (idée #9) : parser `chapters` de `dump-single-json` déjà dispo.
- [ ] `Filtres recherche` (idée #1) : `duration/date/type` mappés sur args yt-dlp + filtre local, 0 coût API.
- [ ] `Pagination / infinite scroll` (idée #2) : `page$` déjà câblé, brancher `nextPageToken` scrape via `--flat-playlist --playlist-start`.
### Phase 4 — Option yattee-server sidecar (seulement si Phase 1 insuffisante)
- [ ] `docker-compose/yattee-server.yml` : image `yattee/yattee-server`, `INVIDIOUS_INSTANCE_URL` optionnelle, `ADMIN_USERNAME/PASSWORD` via `.env`.
- [ ] Adaptateur `server/providers/youtube-yattee.mjs` : `GET ${YATTEE_URL}/api/v1/search?q=&type=video` + Basic Auth → `Suggestion[]`. Activé par `YT_SEARCH_MODE=yattee`.
- [ ] Critère GO/NO-GO : si scrape direct tient >95% succès sur 7 j, abandonner sidecar. Sinon le garder pour `trending/comments/captions`.
### Risques / ToS
- Scraping/InnerTube = gris vis-à-vis ToS YouTube, casses fréquentes → prévoir fallback API + version yt-dlp pinnée + alertes `/healthz`.
- Pas de magie IP : datacenter OVH/Hetzner = ban plus vite que résidentiel → prévoir proxy sortant dès le déploiement public.
- Ne jamais logger clés, cookies, PO-Token.
### Critères d'acceptation Step 17
- [x] Sans aucune `YOUTUBE_API_KEY`, `GET /api/search?q=test&providers=yt` retourne `groups.yt[]` (via scrape) et `npm run test:search-e2e` passe. (vérifié live : scrape-only retourne 3 résultats avec durée/vues/channelId)
- [x] Avec quota épuisé simulé (400/403 mock), fallback scrape prend le relais sans 500. (dispatcher scrape-first → api en fallback sur bot-check/timeout/upstream)
- [x] 2ᵉ appel identique < 50 ms (hit cache mémoire/SQLite, pas d'appel yt-dlp). (mesuré : 3971 ms → 30 ms)
- [x] `/healthz` expose `ytdlpVersion`, `mode`, `cacheHitRate`. (`/healthz` + `/api/healthz` : mode, ytdlp bin/version, antiban, cache mem+sqlite, metrics jour, clés)
- [x] Aucune régression `dm/tw/pt/od/ru`. (`test:search-e2e` + `test:suggest` verts)
### Implémenté le 2026-09-25 (Step 17 DONE)
- Nouveaux : `server/providers/youtube-common.mjs` (clés, `YT_SEARCH_MODE`, args anti-ban cookies/PO-Token/proxy, hash, métriques), `server/providers/youtube-scrape.mjs` (search/channel/trending via `yt-dlp --flat-playlist`, parsers `mapFlatEntry`/`parseFlatPlaylistJson` testés offline), `db/migrations/20250926_add_youtube_scrape_cache.sql`, `server/tests/youtube-scrape.test.mjs` (`npm run test:ytscrape` + CI).
- Modifiés : `youtube.mjs` (dispatcher scrape-first/api-first/scrape-only/api-only + cache mémoire LRU + SQLite + logs source/latency), `channel-content.mjs` (resolve + contenu YT en scrape-first, fallback API), `index.mjs` (`/healthz`, `/api/trending`, import common), `db.mjs` (helpers cache/métriques jour), `.env.example` (nouvelles vars), `package.json` (`test:ytscrape`, `ytdlp:update`).
- Notes : `/feed/trending` retiré côté YouTube (redirect home) → repli `ytsearch` trié vues. yt-dlp local `2026.08.19` : penser `npm run ytdlp:update`. Sidecar yattee-server abandonné (scrape direct >95% : à confirmer sur 7 j).
- Fix 2026-09-25 (`spawn yt-dlp ENOENT`) : `resolveYtDlpBin()` (`youtube-common.mjs`) avec ordre `YT_DLP_PATH` > PATH > bundled `youtube-dl-exec` (cause : yt-dlp seulement présent via shims scoop utilisateur, invisible d'un autre contexte/Docker) ; fallback API sur **tout** échec scrape en `scrape-first` (plus d'erreur brute en UI) ; message actionnable `youtube_no_source` si ni binaire ni clé ; `/healthz` expose `binOk` + binaire résolu.
---
## Step 18 — InnerTube direct façon SmartTube (pagination illimitée + vidéos connexes) ✅
### Contexte
SmartTube (`yuliskov/SmartTube`, 34k stars) ne touche ni la Data API ni les Google Services : il parle à **InnerTube** (`youtubei/v1/*`, clients TV) via `MediaServiceCore` ("Unofficial Java api for YouTube") — search + continuations, watch-next (`/next`), player. D'où 0 quota et scroll infini.
### Implémenté le 2026-09-25
- Dépendance `youtubei.js` **pinnée `18.1.0`** (API interne non documentée → pin + bump manuel).
- Nouveau `server/providers/youtube-innertube.mjs` : session singleton lazy (`YT_INNERTUBE_GL/HL`), `searchViaInnerTube` (**continuations fusionnées** jusqu'à couvrir `page*limit` — corrige le bug "page 2 vide" : 1 page InnerTube ≈ 17-20 items ≠ `limit`), `getRelatedViaInnerTube` (watch-next), mappers purs `mapVideoNode`/`mapLockupView`/`parseViewsText`/`parseDurationLabel`, chaîne de continuations en mémoire + SQLite existant.
- `mapLockupView` : le watch-next moderne renvoie des **LockupView** (`content_id`, `metadata.title.text`, `content_image.image[]`, chaîne + vues dans `metadata_rows`, durée dans le label a11y) — filtrés avant, mappés maintenant.
- Dispatcher `innertube-first` (**défaut**) : InnerTube → scrape yt-dlp → API officielle ; `innertube-only` ajouté ; les anciens modes inchangés ; on ne persiste plus les résultats vides (anti-empoisonnement du cache).
- `GET /api/details/youtube/:videoId` → champ **`related[]`** (24 items, cache mémoire 1h, `?related=0` pour désactiver, best-effort jamais bloquant).
- Tests `server/tests/youtube-innertube.test.mjs` (`npm run test:ytinnertube` + CI), `.env.example` à jour.
- Vérifié live : recherche p1=50 + p2=50 = **100 uniques** (fin de la limite 24), related=24 avec titres FR/durées/vues, `test:search-e2e` + `test:suggest` verts.
- Limites connues : pas de `getTrending` en youtubei.js v18 → trending reste sur scrape ; continuations parfois redondantes (dédup + garde-fou 12) ; même combat anti-ban IP qu'avant (cookies/PO-Token/proxy réutilisés côté yt-dlp, session InnerTube sans auth).
- Complément 2026-09-26 : Watch → sidebar « connexes » branchée sur le vrai watch-next (`loadRelatedSuggestions()` utilise `GET /api/details/youtube/:id` → `related[]`, fallback recherche-par-titre pour les autres providers et en cas d'échec) ; `docker-compose/.env.example` + `README.md` (endpoints + roadmap) à jour ; idée #2 (pagination/infinite scroll) clôturée.
- Complément transcript InnerTube : découverte des pistes via `getInfo().captions` (`YT_TRANSCRIPT_SOURCE`, défaut `innertube-first`) adaptée au format yt-dlp → `pickTrack`/`orderedTracks`/`parseTrackText` inchangés ; 0 piste InnerTube = `no_subtitles` direct (même backend que le lecteur) ; fallback yt-dlp conservé. Vérifié live (`lang=en` → 217 lignes) ; `test:transcript` 17/17.