feat(youtube): InnerTube-first search, transcripts and watch-next related (Steps 15-18)
- InnerTube layer via pinned youtubei.js 18.1.0 (no quota, no key): search with merged continuations (unlimited pages), watch-next related with LockupView mapping, caption-track discovery - 3-layer dispatcher (YT_SEARCH_MODE, default innertube-first): innertube -> yt-dlp scrape -> official API, graceful errors.yt - Robust yt-dlp binary resolution (YT_DLP_PATH > PATH > bundled) with systematic API fallback (fixes spawn ENOENT in UI) - Transcript: InnerTube caption discovery (YT_TRANSCRIPT_SOURCE), reusing pickTrack/orderedTracks/parseTrackText; yt-dlp fallback kept - Watch: sidebar uses real watch-next related[] (/api/details), title-search fallback for other providers - Cache: memory LRU + SQLite (youtube_search_cache, youtube_metrics), never persist empty pages; /healthz observability; /api/trending - Includes pending Step 15/16 leftovers in same files (suggest, test scripts); unrelated provider adapters left uncommitted
This commit is contained in:
+21
-1
@@ -1,8 +1,28 @@
|
|||||||
# Configuration des clés API pour les adaptateurs de recherche
|
# Configuration des clés API pour les adaptateurs de recherche
|
||||||
# Copiez ce fichier en .env et remplissez les valeurs
|
# Copiez ce fichier en .env et remplissez les valeurs
|
||||||
|
|
||||||
# YouTube API Key (obligatoire pour la recherche YouTube)
|
# --- YouTube hybride Step 17-18 : InnerTube (sans clé) > scrape yt-dlp > API officielle ---
|
||||||
|
# Mode : innertube-first (défaut, 0 quota, pagination illimitée) | innertube-only
|
||||||
|
# | scrape-first | api-first | scrape-only | api-only
|
||||||
|
YT_SEARCH_MODE=innertube-first
|
||||||
|
# Localisation InnerTube (gl=region, hl=langue)
|
||||||
|
YT_INNERTUBE_GL=FR
|
||||||
|
YT_INNERTUBE_HL=fr
|
||||||
|
# Découverte des pistes de sous-titres YouTube : innertube-first (défaut,
|
||||||
|
# sans spawn yt-dlp) | ytdlp-only | innertube-only
|
||||||
|
YT_TRANSCRIPT_SOURCE=innertube-first
|
||||||
|
# Cache scrape (30 min) + binaire yt-dlp (frais requis contre le bot-check)
|
||||||
|
YT_SCRAPE_TTL_MS=1800000
|
||||||
|
YT_DLP_TIMEOUT_MS=20000
|
||||||
|
# YT_DLP_PATH=/usr/local/bin/yt-dlp
|
||||||
|
# Anti-ban : cookies exportés (volume :ro en Docker) + PO-Token + proxy sortant
|
||||||
|
# YT_COOKIES_FILE=/cookies/youtube.txt
|
||||||
|
# YT_PO_TOKEN=
|
||||||
|
# YT_EGRESS_PROXY=http://proxy:8080
|
||||||
|
|
||||||
|
# YouTube API Key (fallback uniquement en scrape-first, obligatoire en api-only)
|
||||||
# Obtenez une clé ici : https://console.developers.google.com/
|
# Obtenez une clé ici : https://console.developers.google.com/
|
||||||
|
# Accepte YOUTUBE_API_KEYS en CSV ('k1,k2') ou JSON ('["k1","k2"]') avec rotation auto.
|
||||||
YOUTUBE_API_KEY=your_youtube_api_key_here
|
YOUTUBE_API_KEY=your_youtube_api_key_here
|
||||||
|
|
||||||
# Twitch Client ID (optionnel, pour une recherche Twitch plus complète)
|
# Twitch Client ID (optionnel, pour une recherche Twitch plus complète)
|
||||||
|
|||||||
@@ -42,3 +42,15 @@ jobs:
|
|||||||
|
|
||||||
- name: E2E scenarios (search UX)
|
- name: E2E scenarios (search UX)
|
||||||
run: npm run test:search-e2e
|
run: npm run test:search-e2e
|
||||||
|
|
||||||
|
- name: Suggest typeahead (Step 15)
|
||||||
|
run: npm run test:suggest
|
||||||
|
|
||||||
|
- name: Transcripts (Step 16)
|
||||||
|
run: npm run test:transcript
|
||||||
|
|
||||||
|
- name: YouTube scrape-first (Step 17, offline)
|
||||||
|
run: npm run test:ytscrape
|
||||||
|
|
||||||
|
- name: YouTube InnerTube (Step 18, offline)
|
||||||
|
run: npm run test:ytinnertube
|
||||||
|
|||||||
@@ -56,20 +56,28 @@ Un seul champ, tous les fournisseurs — avec filtres, raccourcis et deep-links
|
|||||||
* **Deep-links** : `/#/search?q=…&providers=yt,ru` relance la recherche filtrée — partageable
|
* **Deep-links** : `/#/search?q=…&providers=yt,ru` relance la recherche filtrée — partageable
|
||||||
* **Fallback préférence** : URL sans `providers` → préférence `defaultProviders` de l’utilisateur → provider actif
|
* **Fallback préférence** : URL sans `providers` → préférence `defaultProviders` de l’utilisateur → provider actif
|
||||||
* **Accessibilité** : focus trap dans les modals, Esc pour fermer, aria-combobox sur le champ, focus restauré à la fermeture
|
* **Accessibilité** : focus trap dans les modals, Esc pour fermer, aria-combobox sur le champ, focus restauré à la fermeture
|
||||||
|
* **Typeahead requête** : sous l’input, suggestions de requêtes (recherches récentes 🕘 + groupes par provider `YT/DM/…`, sous-chaîne surlignée) — `GET /api/search/suggest?q=…&providers=…&limit=…` (min 2 caractères, debounce 250 ms, cache 5 min, dégradation `[]` par provider) ; ↑/↓/Enter/Tab/Esc, priorité au popover `@`
|
||||||
|
|
||||||
<!-- TODO: add docs/search-ux.gif (capture des chips, du picker Ctrl+K et du deep-link) -->
|
<!-- TODO: add docs/search-ux.gif (capture des chips, du picker Ctrl+K et du deep-link) -->
|
||||||
|
|
||||||
### Endpoints API concernés
|
### Endpoints API concernés
|
||||||
|
|
||||||
* `GET /api/search?q=…&providers=yt,dm` — fan-out parallèle, réponse groupée par provider
|
* `GET /api/search?q=…&providers=yt,dm` — fan-out parallèle, réponse groupée par provider (`page`/`pageSize`/`sort` ; YT sans quota via InnerTube + continuations)
|
||||||
|
* `GET /api/search/suggest?q=…&providers=yt,dm&limit=10` — typeahead `{ q, groups: { yt: string[], … } }` (cache 5 min, rate-limit)
|
||||||
|
* `GET /api/details/youtube/:videoId` — métadonnées + `related[]` (watch-next InnerTube, `?related=0` pour désactiver)
|
||||||
|
* `GET /api/trending?provider=yt&limit=…` — tendances YT sans clé
|
||||||
|
* `GET /healthz` (alias `/api/healthz`) — mode YT, binaire yt-dlp `binOk`, cache, métriques quota/jour, clés
|
||||||
|
* `GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=` — transcript `{ lang, available, languages, lines: [{ t, dur, text }] }` (cache 24 h, rate-limit 10/min ; absent → 200 `{ available: false }`, échec → 502 ; YouTube : découverte des pistes via InnerTube, `YT_TRANSCRIPT_SOURCE`)
|
||||||
* `GET/PATCH /api/user/preferences` — `defaultProviders` (tableau JSON, sanitizé serveur)
|
* `GET/PATCH /api/user/preferences` — `defaultProviders` (tableau JSON, sanitizé serveur)
|
||||||
* `POST /api/telemetry/events` — événements UX anonymes (whitelist : `search_submit`, `provider_picker_open`, `provider_apply`, `at_autocomplete_use`, `quick_menu_open`)
|
* `POST /api/telemetry/events` — événements UX anonymes (whitelist : `search_submit`, `provider_picker_open`, `provider_apply`, `at_autocomplete_use`, `quick_menu_open`, `suggest_shown`, `suggest_used`)
|
||||||
|
|
||||||
### Tests
|
### Tests
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
npm run test:search # unitaires SearchService + parsing @ + picker
|
npm run test:search # unitaires SearchService + parsing @ + picker
|
||||||
npm run test:search-e2e # scénarios e2e (serveur réel isolé)
|
npm run test:search-e2e # scénarios e2e (serveur réel isolé)
|
||||||
|
npm run test:suggest # typeahead : parsing/dédup + contrat /api/search/suggest
|
||||||
|
npm run test:transcript # transcripts : parseurs json3/vtt + contrat /api/transcript
|
||||||
npm run test:preferences # persistance defaultProviders
|
npm run test:preferences # persistance defaultProviders
|
||||||
npm run test:telemetry # télémétrie minimale
|
npm run test:telemetry # télémétrie minimale
|
||||||
```
|
```
|
||||||
@@ -242,7 +250,8 @@ Ajoutez au besoin `log-driver`, `log-opts`, `default-address-pools`, etc.
|
|||||||
* ✅ **Téléchargements** intégrés — file d'attente **persistée en SQLite** (survit aux redémarrages API), jobs **par utilisateur** (répertoires isolés, ownership sur status/fichier/cancel), **quota de stockage** configurable (`DOWNLOAD_STORAGE_QUOTA_BYTES`, fenêtre `DOWNLOAD_QUOTA_WINDOW_MS`), **reprise** des jobs échoués/interrrompus (bouton Réessayer), **page Bibliothèque > Téléchargements** (filtres par état, progression live, quota), nettoyage auto des fichiers orphelins au boot
|
* ✅ **Téléchargements** intégrés — file d'attente **persistée en SQLite** (survit aux redémarrages API), jobs **par utilisateur** (répertoires isolés, ownership sur status/fichier/cancel), **quota de stockage** configurable (`DOWNLOAD_STORAGE_QUOTA_BYTES`, fenêtre `DOWNLOAD_QUOTA_WINDOW_MS`), **reprise** des jobs échoués/interrrompus (bouton Réessayer), **page Bibliothèque > Téléchargements** (filtres par état, progression live, quota), nettoyage auto des fichiers orphelins au boot
|
||||||
* 🔧 Variables : `DOWNLOAD_MAX_CONCURRENT` (2), `DOWNLOAD_STORAGE_QUOTA_BYTES` (5 GiB), `DOWNLOAD_QUOTA_WINDOW_MS` (30 j), `DOWNLOAD_PROVIDERS` (peertube,odysee)
|
* 🔧 Variables : `DOWNLOAD_MAX_CONCURRENT` (2), `DOWNLOAD_STORAGE_QUOTA_BYTES` (5 GiB), `DOWNLOAD_QUOTA_WINDOW_MS` (30 j), `DOWNLOAD_PROVIDERS` (peertube,odysee)
|
||||||
* ⏳ **Import/Export** playlists (JSON / OPML-like)
|
* ⏳ **Import/Export** playlists (JSON / OPML-like)
|
||||||
* ⏳ **Sous-titres & transcripts** (si dispo API; fallback parsing)
|
* ✅ **Sous-titres & transcripts** — `GET /api/transcript/:provider/:videoId` (yt-dlp `subtitles`/`automatic_captions`, parsing `json3`/`vtt`, cache 24 h, rate-limit), panneau **Transcript** sur la page Watch (sélecteur de langue, dégradation propre si indisponible)
|
||||||
|
* ✅ **YouTube sans quota (InnerTube façon SmartTube)** — `youtubei.js` pinné, chaîne `innertube → scrape yt-dlp → API officielle` (`YT_SEARCH_MODE`, défaut `innertube-first`), **pagination illimitée** via continuations (scroll infini, `page=2,3…`), **vidéos connexes** watch-next dans `GET /api/details/youtube/:videoId` → sidebar Watch, `GET /api/trending?provider=yt`, cache mémoire + SQLite, anti-ban (`YT_COOKIES_FILE`, `YT_PO_TOKEN`, `YT_EGRESS_PROXY`), observabilité `/healthz` (mode, binaire `binOk`, cache, quota jour, clés)
|
||||||
* ⏳ **PWA** (installable, offline cache des métadonnées)
|
* ⏳ **PWA** (installable, offline cache des métadonnées)
|
||||||
* ⏳ **Chromecast / AirPlay**
|
* ⏳ **Chromecast / AirPlay**
|
||||||
* ⏳ Mode “TV”
|
* ⏳ Mode “TV”
|
||||||
|
|||||||
@@ -0,0 +1,18 @@
|
|||||||
|
-- Step 17 : cache persistant recherche YouTube (scrape + API) + métriques quota journalières.
|
||||||
|
CREATE TABLE IF NOT EXISTS youtube_search_cache (
|
||||||
|
q_hash TEXT PRIMARY KEY,
|
||||||
|
q TEXT NOT NULL,
|
||||||
|
payload_json TEXT NOT NULL,
|
||||||
|
source TEXT NOT NULL DEFAULT 'scrape',
|
||||||
|
created_at INTEGER NOT NULL,
|
||||||
|
expires_at INTEGER NOT NULL
|
||||||
|
);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_yt_cache_exp ON youtube_search_cache(expires_at);
|
||||||
|
|
||||||
|
CREATE TABLE IF NOT EXISTS youtube_metrics (
|
||||||
|
day TEXT PRIMARY KEY,
|
||||||
|
scrape_calls INTEGER NOT NULL DEFAULT 0,
|
||||||
|
api_calls INTEGER NOT NULL DEFAULT 0,
|
||||||
|
quota_units INTEGER NOT NULL DEFAULT 0,
|
||||||
|
updated_at TEXT NOT NULL
|
||||||
|
);
|
||||||
@@ -15,6 +15,23 @@ TWITCH_CLIENT_SECRET=votre_client_secret_twitch_ici
|
|||||||
# Configuration du cache (en millisecondes)
|
# Configuration du cache (en millisecondes)
|
||||||
YT_CACHE_TTL_MS=3600000 # 1 heure par défaut
|
YT_CACHE_TTL_MS=3600000 # 1 heure par défaut
|
||||||
|
|
||||||
|
# --- YouTube hybride : InnerTube (sans clé) > scrape yt-dlp > API officielle ---
|
||||||
|
# Mode : innertube-first (défaut, 0 quota, pagination illimitée) | innertube-only
|
||||||
|
# | scrape-first | api-first | scrape-only | api-only
|
||||||
|
YT_SEARCH_MODE=innertube-first
|
||||||
|
YT_INNERTUBE_GL=FR
|
||||||
|
YT_INNERTUBE_HL=fr
|
||||||
|
# Sous-titres YouTube : innertube-first (défaut) | ytdlp-only | innertube-only
|
||||||
|
YT_TRANSCRIPT_SOURCE=innertube-first
|
||||||
|
# Cache scrape/pagination (30 min) + timeout yt-dlp
|
||||||
|
YT_SCRAPE_TTL_MS=1800000
|
||||||
|
YT_DLP_TIMEOUT_MS=20000
|
||||||
|
# Anti-ban (optionnel) : binaire explicite, cookies exportés, PO-Token, proxy sortant
|
||||||
|
# YT_DLP_PATH=/usr/local/bin/yt-dlp
|
||||||
|
# YT_COOKIES_FILE=/cookies/youtube.txt
|
||||||
|
# YT_PO_TOKEN=
|
||||||
|
# YT_EGRESS_PROXY=http://proxy:8080
|
||||||
|
|
||||||
# Téléchargements : providers autorisés (liste CSV) + concurrence max par user
|
# Téléchargements : providers autorisés (liste CSV) + concurrence max par user
|
||||||
DOWNLOAD_PROVIDERS=youtube,dailymotion,twitch,peertube,odysee,rumble
|
DOWNLOAD_PROVIDERS=youtube,dailymotion,twitch,peertube,odysee,rumble
|
||||||
DOWNLOAD_MAX_CONCURRENT=2
|
DOWNLOAD_MAX_CONCURRENT=2
|
||||||
|
|||||||
@@ -0,0 +1,769 @@
|
|||||||
|
# Guide d'architecture et de développement — Transcripts multi-fournisseurs pour NewTube
|
||||||
|
|
||||||
|
**Version :** 1.0
|
||||||
|
**Date :** 2026-09-25
|
||||||
|
**Statut :** Proposition d'implémentation
|
||||||
|
**Périmètre :** Ajout de la fonctionnalité « Transcripts / Sous-titres » sur NewTube, pour les 6 fournisseurs supportés.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Table des matières
|
||||||
|
|
||||||
|
1. [Contexte](#1-contexte)
|
||||||
|
2. [Objectifs et non-objectifs](#2-objectifs-et-non-objectifs)
|
||||||
|
3. [Architecture globale](#3-architecture-globale)
|
||||||
|
4. [Backend](#4-backend)
|
||||||
|
5. [Frontend](#5-frontend)
|
||||||
|
6. [Référence API](#6-référence-api)
|
||||||
|
7. [Tests](#7-tests)
|
||||||
|
8. [Développement pas à pas](#8-développement-pas-à-pas)
|
||||||
|
9. [Déploiement et configuration](#9-déploiement-et-configuration)
|
||||||
|
10. [Risques et mitigations](#10-risques-et-mitigations)
|
||||||
|
11. [Alternatives écartées](#11-alternatives-écartées)
|
||||||
|
12. [Roadmap](#12-roadmap)
|
||||||
|
13. [Annexes](#13-annexes)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Contexte
|
||||||
|
|
||||||
|
NewTube est un agrégateur multi-fournisseurs qui s'appuie déjà sur `yt-dlp` via `youtube-dl-exec` pour extraire les métadonnées et les formats de téléchargement.
|
||||||
|
|
||||||
|
Fournisseurs supportés :
|
||||||
|
|
||||||
|
- YouTube
|
||||||
|
- Dailymotion
|
||||||
|
- Twitch
|
||||||
|
- PeerTube
|
||||||
|
- Odysee
|
||||||
|
- Rumble
|
||||||
|
|
||||||
|
L'analyse du code existant montre que :
|
||||||
|
|
||||||
|
- `server/index.mjs:10` importe `youtube-dl-exec`.
|
||||||
|
- `providerUrlFrom()` (`server/index.mjs:541`) construit une URL normalisée pour les 6 fournisseurs.
|
||||||
|
- Les routes `/api/details/:provider/:videoId` (`l.988`) et `/api/download/.../formats` (`l.1061`) utilisent déjà `youtubedl(url, { dumpSingleJson, skipDownload })`.
|
||||||
|
- `yt-dlp` renvoie dans ce même JSON les champs `subtitles` et `automatic_captions`.
|
||||||
|
- Côté UI, `watch.component.html:154` et `:165` fournissent le motif du panneau « Download » à copier pour un panneau « Transcript ».
|
||||||
|
- Le README contient déjà une ligne roadmap : `⏳ Sous-titres & transcripts` (`ligne 245`).
|
||||||
|
|
||||||
|
La fonctionnalité peut donc être ajoutée **sans nouveau registre de providers**, **sans adaptateur par fournisseur**, et **sans migration de base de données**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Objectifs et non-objectifs
|
||||||
|
|
||||||
|
### 2.1 Objectifs
|
||||||
|
|
||||||
|
- Extraire les sous-titres manuels et automatiques via `yt-dlp`.
|
||||||
|
- Supporter plusieurs langues avec un sélecteur.
|
||||||
|
- Offrir une dégradation propre quand aucune piste n'existe.
|
||||||
|
- Mettre en cache les transcripts (immuables) pour éviter les appels répétés.
|
||||||
|
- Limiter le débit pour éviter les erreurs HTTP 429 de YouTube.
|
||||||
|
- Ajouter une UI simple dans la page « Watch ».
|
||||||
|
- Rester compatible avec les 6 fournisseurs sans code spécifique par plateforme.
|
||||||
|
|
||||||
|
### 2.2 Non-objectifs
|
||||||
|
|
||||||
|
- Traduction automatique des transcripts.
|
||||||
|
- Transcription audio via LLM ou service tiers.
|
||||||
|
- Téléchargement de la vidéo complète.
|
||||||
|
- Authentification OAuth.
|
||||||
|
- Seek avancé dans la vidéo depuis le transcript (optionnel, phase 2).
|
||||||
|
- Support garanti des sous-titres sur Twitch, Odysee et Rumble.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Architecture globale
|
||||||
|
|
||||||
|
### 3.1 Schéma
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
A[Client Watch] -->|GET /api/transcript/:provider/:videoId| B[API NewTube]
|
||||||
|
B --> C{Cache ?}
|
||||||
|
C -->|Oui| D[Retour JSON]
|
||||||
|
C -->|Non| E[providerUrlFrom]
|
||||||
|
E --> F[yt-dlp dumpSingleJson]
|
||||||
|
F --> G[subtitles / automatic_captions]
|
||||||
|
G --> H[pickTrack]
|
||||||
|
H --> I[Fetch piste json3/vtt]
|
||||||
|
I --> J[parseJson3 / parseVtt]
|
||||||
|
J --> K[Mise en cache]
|
||||||
|
K --> D
|
||||||
|
D --> A
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 Composants
|
||||||
|
|
||||||
|
| Composant | Fichier | Rôle |
|
||||||
|
|---|---|---|
|
||||||
|
| Module transcript | `server/transcript.mjs` | Fonctions pures : sélection de piste, parsing json3/vtt |
|
||||||
|
| Route API | `server/index.mjs` | Endpoint `/api/transcript/...`, cache, rate-limit |
|
||||||
|
| UI Watch | `watch.component.ts/html` | Bouton, panneau, sélecteur de langue, affichage |
|
||||||
|
| Tests | `server/tests/transcript.test.mjs` | Tests unitaires et d'intégration |
|
||||||
|
| Script npm | `package.json` | `test:transcript` |
|
||||||
|
| Documentation | `README.md` | Case roadmap cochée + bump de version |
|
||||||
|
|
||||||
|
### 3.3 Principe directeur
|
||||||
|
|
||||||
|
> **Rung 2 : réutiliser, pas réinventer.**
|
||||||
|
|
||||||
|
Une seule route, un seul module de parsing, une seule UI. La capacité dépend de la plateforme, mais le code ne change pas.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Backend
|
||||||
|
|
||||||
|
### 4.1 Module `server/transcript.mjs`
|
||||||
|
|
||||||
|
Fonctions pures, testables sans réseau ni base de données.
|
||||||
|
|
||||||
|
#### 4.1.1 `pickTrack(json, lang)`
|
||||||
|
|
||||||
|
Sélectionne la meilleure piste disponible.
|
||||||
|
|
||||||
|
Priorités :
|
||||||
|
|
||||||
|
1. Sous-titres manuels (`subtitles`) dans la langue demandée.
|
||||||
|
2. Sous-titres manuels dans une langue proche (`fr-*`).
|
||||||
|
3. Sous-titres automatiques (`automatic_captions`) dans la langue demandée.
|
||||||
|
4. Sous-titres automatiques dans une langue proche.
|
||||||
|
5. Première piste disponible.
|
||||||
|
|
||||||
|
```js
|
||||||
|
export function pickTrack(json, lang = 'fr') {
|
||||||
|
const manual = json.subtitles || {};
|
||||||
|
const auto = json.automatic_captions || {};
|
||||||
|
const all = { ...auto, ...manual };
|
||||||
|
const languages = Object.keys(all);
|
||||||
|
|
||||||
|
if (!languages.length) {
|
||||||
|
return { track: null, languages: [], lang: null };
|
||||||
|
}
|
||||||
|
|
||||||
|
const findLang = (dict, code) =>
|
||||||
|
dict[code] || Object.keys(dict).find(k => k.startsWith(code + '-'));
|
||||||
|
|
||||||
|
const chosenLang =
|
||||||
|
findLang(manual, lang) ? lang :
|
||||||
|
findLang(auto, lang) ? lang :
|
||||||
|
languages[0];
|
||||||
|
|
||||||
|
const track =
|
||||||
|
manual[chosenLang] ||
|
||||||
|
auto[chosenLang] ||
|
||||||
|
manual[languages[0]] ||
|
||||||
|
auto[languages[0]];
|
||||||
|
|
||||||
|
return { track, languages, lang: chosenLang };
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 4.1.2 `parseJson3(data)`
|
||||||
|
|
||||||
|
Convertit le format `json3` de YouTube en lignes normalisées.
|
||||||
|
|
||||||
|
```js
|
||||||
|
export function parseJson3(data) {
|
||||||
|
const events = data.events || [];
|
||||||
|
return events
|
||||||
|
.filter(e => e.segs)
|
||||||
|
.map(e => ({
|
||||||
|
t: (e.tStartMs || 0) / 1000,
|
||||||
|
dur: (e.dDurationMs || 0) / 1000,
|
||||||
|
text: e.segs.map(s => s.utf8).join('').trim()
|
||||||
|
}))
|
||||||
|
.filter(line => line.text);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 4.1.3 `parseVtt(text)`
|
||||||
|
|
||||||
|
Convertit un fichier VTT en lignes normalisées.
|
||||||
|
|
||||||
|
```js
|
||||||
|
export function parseVtt(text) {
|
||||||
|
const lines = [];
|
||||||
|
const blocks = text.split(/\n\n+/);
|
||||||
|
|
||||||
|
for (const block of blocks) {
|
||||||
|
const match = block.match(/(\d{2}):(\d{2}):(\d{2})\.(\d{3}) --> (\d{2}):(\d{2}):(\d{2})\.(\d{3})/);
|
||||||
|
if (!match) continue;
|
||||||
|
|
||||||
|
const start =
|
||||||
|
parseInt(match[1]) * 3600 +
|
||||||
|
parseInt(match[2]) * 60 +
|
||||||
|
parseInt(match[3]) +
|
||||||
|
parseInt(match[4]) / 1000;
|
||||||
|
|
||||||
|
const end =
|
||||||
|
parseInt(match[5]) * 3600 +
|
||||||
|
parseInt(match[6]) * 60 +
|
||||||
|
parseInt(match[7]) +
|
||||||
|
parseInt(match[8]) / 1000;
|
||||||
|
|
||||||
|
const content = block
|
||||||
|
.split('\n')
|
||||||
|
.filter(line => !line.includes('-->') && !/^\d+$/.test(line))
|
||||||
|
.join(' ')
|
||||||
|
.trim();
|
||||||
|
|
||||||
|
if (content) {
|
||||||
|
lines.push({ t: start, dur: end - start, text: content });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return lines;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 4.1.4 Structure de retour normalisée
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"lang": "fr",
|
||||||
|
"available": true,
|
||||||
|
"languages": ["fr", "en", "es"],
|
||||||
|
"lines": [
|
||||||
|
{ "t": 0, "dur": 2.5, "text": "Bonjour à tous" },
|
||||||
|
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue dans cette vidéo" }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
En l'absence de piste :
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"lang": null,
|
||||||
|
"available": false,
|
||||||
|
"languages": [],
|
||||||
|
"lines": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 4.2 Route API dans `server/index.mjs`
|
||||||
|
|
||||||
|
#### 4.2.1 Signature
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 4.2.2 Flux d'exécution
|
||||||
|
|
||||||
|
1. Valider le provider.
|
||||||
|
2. Construire la clé de cache : `transcript:${provider}:${videoId}:${lang}`.
|
||||||
|
3. Vérifier le cache.
|
||||||
|
4. Si absent :
|
||||||
|
- `providerUrlFrom()` construit l'URL.
|
||||||
|
- `youtubedl(url, { dumpSingleJson: true, skipDownload: true })`.
|
||||||
|
- `pickTrack(json, lang)`.
|
||||||
|
- Si aucune piste : répondre `{ available: false }` avec HTTP 200.
|
||||||
|
- Sinon : `fetch(track.url)`.
|
||||||
|
- Parser selon le format (`json3` ou `vtt`).
|
||||||
|
- Mettre en cache.
|
||||||
|
5. Répondre.
|
||||||
|
|
||||||
|
#### 4.2.3 Pseudo-code
|
||||||
|
|
||||||
|
```js
|
||||||
|
app.get('/api/transcript/:provider/:videoId', channelsLimiter, async (req, res) => {
|
||||||
|
const { provider, videoId } = req.params;
|
||||||
|
const { lang = 'fr', instance, slug, sourceUrl } = req.query;
|
||||||
|
|
||||||
|
const cacheKey = `transcript:${provider}:${videoId}:${lang}`;
|
||||||
|
const cached = transcriptCache.get(cacheKey);
|
||||||
|
if (cached) return res.json(cached);
|
||||||
|
|
||||||
|
try {
|
||||||
|
const url = providerUrlFrom(provider, videoId, { instance, slug, sourceUrl });
|
||||||
|
const json = await youtubedl(url, {
|
||||||
|
dumpSingleJson: true,
|
||||||
|
skipDownload: true,
|
||||||
|
});
|
||||||
|
|
||||||
|
const { track, languages, lang: chosenLang } = pickTrack(json, lang);
|
||||||
|
|
||||||
|
if (!track) {
|
||||||
|
const empty = { lang: null, available: false, languages: [], lines: [] };
|
||||||
|
transcriptCache.set(cacheKey, empty, TTL_LONG);
|
||||||
|
return res.json(empty);
|
||||||
|
}
|
||||||
|
|
||||||
|
const response = await fetch(track.url);
|
||||||
|
const text = await response.text();
|
||||||
|
|
||||||
|
const lines = track.ext === 'json3'
|
||||||
|
? parseJson3(JSON.parse(text))
|
||||||
|
: parseVtt(text);
|
||||||
|
|
||||||
|
const result = {
|
||||||
|
lang: chosenLang,
|
||||||
|
available: true,
|
||||||
|
languages,
|
||||||
|
lines,
|
||||||
|
};
|
||||||
|
|
||||||
|
transcriptCache.set(cacheKey, result, TTL_LONG);
|
||||||
|
res.json(result);
|
||||||
|
} catch (err) {
|
||||||
|
console.error('Transcript error:', err);
|
||||||
|
res.status(502).json({ available: false, error: 'transcript_fetch_failed' });
|
||||||
|
}
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 4.3 Cache
|
||||||
|
|
||||||
|
- **Type :** `Map` avec TTL, comme `ytCache` (`server/index.mjs:838`).
|
||||||
|
- **Clé :** `transcript:${provider}:${videoId}:${lang}`.
|
||||||
|
- **TTL :** long (ex. 24 h) car les transcripts sont immuables.
|
||||||
|
- **Évolution :** pour un déploiement multi-instances, remplacer par Redis ou Memcached.
|
||||||
|
|
||||||
|
```js
|
||||||
|
const transcriptCache = {
|
||||||
|
data: new Map(),
|
||||||
|
get(key) {
|
||||||
|
const entry = this.data.get(key);
|
||||||
|
if (!entry) return null;
|
||||||
|
if (Date.now() > entry.expires) {
|
||||||
|
this.data.delete(key);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
return entry.value;
|
||||||
|
},
|
||||||
|
set(key, value, ttlMs) {
|
||||||
|
this.data.set(key, { value, expires: Date.now() + ttlMs });
|
||||||
|
}
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 4.4 Rate limiting
|
||||||
|
|
||||||
|
- Réutiliser le motif `channelsLimiter`.
|
||||||
|
- Limiter par IP, par exemple 10 requêtes / minute sur `/api/transcript/...`.
|
||||||
|
- En cas de HTTP 429 de YouTube :
|
||||||
|
- Ne pas réessayer immédiatement.
|
||||||
|
- Attendre 20 s minimum.
|
||||||
|
- Utiliser un backoff exponentiel.
|
||||||
|
- Si le problème persiste, activer le fallback.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 4.5 Fallback en cas de 429 persistant
|
||||||
|
|
||||||
|
Si le endpoint `timedtext` de YouTube renvoie trop de 429 :
|
||||||
|
|
||||||
|
1. Laisser `yt-dlp` écrire le VTT dans un dossier temporaire.
|
||||||
|
2. Lire le fichier.
|
||||||
|
3. Parser avec `parseVtt()`.
|
||||||
|
|
||||||
|
Pattern déjà utilisé pour les downloads : `youtubedl.exec`.
|
||||||
|
|
||||||
|
```js
|
||||||
|
const tmp = path.join(os.tmpdir(), `transcript-${Date.now()}.vtt`);
|
||||||
|
await youtubedl.exec(url, {
|
||||||
|
writeAutoSub: true,
|
||||||
|
subLang: lang,
|
||||||
|
subFormat: 'vtt',
|
||||||
|
output: tmp,
|
||||||
|
});
|
||||||
|
const text = await fs.readFile(tmp, 'utf8');
|
||||||
|
const lines = parseVtt(text);
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Frontend
|
||||||
|
|
||||||
|
### 5.1 UI dans `watch.component`
|
||||||
|
|
||||||
|
#### 5.1.1 Bouton
|
||||||
|
|
||||||
|
Ajouter un bouton « Transcript » à côté du bouton « Download ».
|
||||||
|
|
||||||
|
```html
|
||||||
|
<button class="btn-transcript" (click)="toggleTranscript()">
|
||||||
|
Transcript
|
||||||
|
</button>
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 5.1.2 Panneau
|
||||||
|
|
||||||
|
Copier le motif du panneau Download (`watch.component.html:165`).
|
||||||
|
|
||||||
|
```html
|
||||||
|
<div *ngIf="transcriptOpen" class="transcript-panel">
|
||||||
|
<div class="transcript-header">
|
||||||
|
<h3>Transcript</h3>
|
||||||
|
<select [(ngModel)]="selectedLang" (change)="loadTranscript()">
|
||||||
|
<option *ngFor="let lang of transcriptLanguages" [value]="lang">
|
||||||
|
{{ lang }}
|
||||||
|
</option>
|
||||||
|
</select>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div *ngIf="transcriptLoading" class="transcript-loading">
|
||||||
|
Chargement…
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div *ngIf="!transcriptLoading && !transcriptAvailable" class="transcript-empty">
|
||||||
|
Aucun sous-titre disponible pour cette vidéo.
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div *ngIf="transcriptAvailable" class="transcript-lines">
|
||||||
|
<div *ngFor="let line of transcriptLines" class="transcript-line">
|
||||||
|
<span class="transcript-time">{{ line.t | duration }}</span>
|
||||||
|
<span class="transcript-text">{{ line.text }}</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 5.1.3 Composant TypeScript
|
||||||
|
|
||||||
|
```ts
|
||||||
|
transcriptOpen = false;
|
||||||
|
transcriptLoading = false;
|
||||||
|
transcriptAvailable = false;
|
||||||
|
transcriptLanguages: string[] = [];
|
||||||
|
transcriptLines: { t: number; dur: number; text: string }[] = [];
|
||||||
|
selectedLang = 'fr';
|
||||||
|
|
||||||
|
toggleTranscript() {
|
||||||
|
this.transcriptOpen = !this.transcriptOpen;
|
||||||
|
if (this.transcriptOpen && !this.transcriptLines.length) {
|
||||||
|
this.loadTranscript();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async loadTranscript() {
|
||||||
|
this.transcriptLoading = true;
|
||||||
|
const res = await fetch(
|
||||||
|
`/api/transcript/${this.provider}/${this.videoId}?lang=${this.selectedLang}`
|
||||||
|
);
|
||||||
|
const data = await res.json();
|
||||||
|
|
||||||
|
this.transcriptAvailable = data.available;
|
||||||
|
this.transcriptLanguages = data.languages || [];
|
||||||
|
this.transcriptLines = data.lines || [];
|
||||||
|
this.transcriptLoading = false;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 5.1.4 Gestion de l'absence de sous-titres
|
||||||
|
|
||||||
|
- Si `available: false`, afficher un message clair.
|
||||||
|
- Le bouton peut rester visible mais le panneau indique l'absence.
|
||||||
|
- Ne jamais casser la page Watch.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 5.2 Variante UI #1 — recommandée
|
||||||
|
|
||||||
|
- Affichage du transcript.
|
||||||
|
- Sélecteur de langue.
|
||||||
|
- Pas de seek.
|
||||||
|
|
||||||
|
**Avantages :** simple, robuste, peu de code.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 5.3 Variante UI #2 — optionnelle
|
||||||
|
|
||||||
|
- Comme #1, plus : clic sur une ligne = seek dans la vidéo.
|
||||||
|
|
||||||
|
**Implémentation :**
|
||||||
|
|
||||||
|
- Player natif : `seekBy` dans `video-player.component.ts:112`.
|
||||||
|
- Iframe YouTube : `seekTo` dans `iframe-progress.service.ts:104`.
|
||||||
|
|
||||||
|
**Inconvénient :** le seek iframe n'est attaché que sous certaines conditions. Il faut brancher par provider. Sensiblement plus de code.
|
||||||
|
|
||||||
|
**Recommandation :** garder #2 pour une phase ultérieure.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Référence API
|
||||||
|
|
||||||
|
### 6.1 Endpoint
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /api/transcript/:provider/:videoId
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.2 Paramètres
|
||||||
|
|
||||||
|
| Nom | Type | Requis | Description |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `provider` | string | Oui | `youtube`, `dailymotion`, `twitch`, `peertube`, `odysee`, `rumble` |
|
||||||
|
| `videoId` | string | Oui | Identifiant de la vidéo |
|
||||||
|
| `lang` | string | Non | Langue souhaitée (ex. `fr`, `en`). Défaut : `fr` |
|
||||||
|
| `instance` | string | Non | Pour PeerTube |
|
||||||
|
| `slug` | string | Non | Pour PeerTube |
|
||||||
|
| `sourceUrl` | string | Non | Pour Odysee / Rumble |
|
||||||
|
|
||||||
|
### 6.3 Réponse succès
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"lang": "fr",
|
||||||
|
"available": true,
|
||||||
|
"languages": ["fr", "en"],
|
||||||
|
"lines": [
|
||||||
|
{ "t": 0, "dur": 2.5, "text": "Bonjour" },
|
||||||
|
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue" }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.4 Réponse absence de sous-titres
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"lang": null,
|
||||||
|
"available": false,
|
||||||
|
"languages": [],
|
||||||
|
"lines": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
HTTP 200.
|
||||||
|
|
||||||
|
### 6.5 Réponse erreur
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"available": false,
|
||||||
|
"error": "transcript_fetch_failed"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
HTTP 502.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Tests
|
||||||
|
|
||||||
|
### 7.1 Tests unitaires
|
||||||
|
|
||||||
|
Fichier : `server/tests/transcript.test.mjs`
|
||||||
|
|
||||||
|
- `pickTrack()` :
|
||||||
|
- manuel prioritaire sur auto.
|
||||||
|
- `fr` → `fr-*` → autre langue.
|
||||||
|
- aucun track → `available: false`.
|
||||||
|
- `parseJson3()` :
|
||||||
|
- events valides.
|
||||||
|
- events sans `segs`.
|
||||||
|
- texte vide filtré.
|
||||||
|
- `parseVtt()` :
|
||||||
|
- blocs valides.
|
||||||
|
- timestamps corrects.
|
||||||
|
- contenu multi-lignes.
|
||||||
|
|
||||||
|
### 7.2 Tests d'intégration
|
||||||
|
|
||||||
|
- Mocker `youtubedl` pour renvoyer un JSON avec `subtitles`.
|
||||||
|
- Vérifier la route `/api/transcript/...`.
|
||||||
|
- Vérifier le cache : deuxième appel ne déclenche pas `youtubedl`.
|
||||||
|
- Vérifier le rate-limit.
|
||||||
|
|
||||||
|
### 7.3 Tests manuels
|
||||||
|
|
||||||
|
| Provider | Vidéo testée | Résultat attendu |
|
||||||
|
|---|---|---|
|
||||||
|
| YouTube | Vidéo avec auto-captions | `available: true`, ~100 langues |
|
||||||
|
| Dailymotion | Vidéo avec pistes | `available: true` |
|
||||||
|
| PeerTube | Vidéo framatube | souvent `available: false` |
|
||||||
|
| Twitch | VOD | `available: false` |
|
||||||
|
| Odysee | Vidéo | `available: false` |
|
||||||
|
| Rumble | Vidéo | `available: false` |
|
||||||
|
|
||||||
|
### 7.4 Script npm
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"scripts": {
|
||||||
|
"test:transcript": "node --test server/tests/transcript.test.mjs"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Développement pas à pas
|
||||||
|
|
||||||
|
### Checklist
|
||||||
|
|
||||||
|
- [ ] Créer `server/transcript.mjs` avec `pickTrack`, `parseJson3`, `parseVtt`.
|
||||||
|
- [ ] Ajouter la route `GET /api/transcript/:provider/:videoId` dans `server/index.mjs`.
|
||||||
|
- [ ] Ajouter le cache `transcriptCache` avec TTL long.
|
||||||
|
- [ ] Ajouter le rate-limit sur le motif `channelsLimiter`.
|
||||||
|
- [ ] Ajouter la gestion des erreurs 429 et le fallback temp dir.
|
||||||
|
- [ ] Créer `server/tests/transcript.test.mjs`.
|
||||||
|
- [ ] Ajouter `test:transcript` dans `package.json`.
|
||||||
|
- [ ] Ajouter le bouton et le panneau dans `watch.component.html`.
|
||||||
|
- [ ] Ajouter la logique dans `watch.component.ts`.
|
||||||
|
- [ ] Mettre à jour le README : cocher `⏳ Sous-titres & transcripts`.
|
||||||
|
- [ ] Bumper la version selon la règle du projet.
|
||||||
|
- [ ] Tester manuellement sur YouTube et Dailymotion.
|
||||||
|
- [ ] Vérifier la dégradation propre sur Twitch, Odysee, Rumble.
|
||||||
|
|
||||||
|
**Volume estimé :** 150 à 200 lignes, 4 fichiers.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Déploiement et configuration
|
||||||
|
|
||||||
|
### 9.1 Variables d'environnement
|
||||||
|
|
||||||
|
| Variable | Description | Défaut |
|
||||||
|
|---|---|---|
|
||||||
|
| `YTDLP_PATH` | Chemin vers `yt-dlp` | `yt-dlp` |
|
||||||
|
| `TRANSCRIPT_CACHE_TTL` | TTL du cache en ms | `86400000` (24 h) |
|
||||||
|
| `TRANSCRIPT_RATE_LIMIT` | Requêtes / minute / IP | `10` |
|
||||||
|
| `TRANSCRIPT_FALLBACK_TMP` | Activer le fallback temp dir | `false` |
|
||||||
|
|
||||||
|
### 9.2 Mise à jour de `yt-dlp`
|
||||||
|
|
||||||
|
- `yt-dlp` évolue vite.
|
||||||
|
- Prévoir une mise à jour régulière.
|
||||||
|
- Surveiller les breaking changes.
|
||||||
|
|
||||||
|
### 9.3 Monitoring
|
||||||
|
|
||||||
|
- Compter les HTTP 429 sur `timedtext`.
|
||||||
|
- Compter les `available: false` par provider.
|
||||||
|
- Mesurer le temps de réponse de `/api/transcript/...`.
|
||||||
|
- Alerter si le taux d'échec dépasse un seuil.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Risques et mitigations
|
||||||
|
|
||||||
|
| Risque | Impact | Mitigation |
|
||||||
|
|---|---|---|
|
||||||
|
| ToS des plateformes | Élevé | Usage personnel / auto-hébergé, ne pas revendre |
|
||||||
|
| HTTP 429 YouTube | Moyen | Cache long, rate-limit, fallback temp dir |
|
||||||
|
| `yt-dlp` cassé | Élevé | Mise à jour régulière, tests de non-régression |
|
||||||
|
| Providers sans sous-titres | Faible | Dégradation propre `available: false` |
|
||||||
|
| Cache mémoire non partagé | Moyen | Redis en production multi-instances |
|
||||||
|
| Seek iframe complexe | Faible | Reporter en phase 2 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Alternatives écartées
|
||||||
|
|
||||||
|
| Alternative | Raison du rejet |
|
||||||
|
|---|---|
|
||||||
|
| YouTube Data API `captions.download` | Nécessite OAuth du propriétaire |
|
||||||
|
| Un implémentation par provider | 6 chemins de code pour le même résultat |
|
||||||
|
| Fetch navigateur `timedtext` | CORS non garanti + casse le proxy clé API |
|
||||||
|
| Services tiers / LLM de transcription | Dépendance, coût, vie privée, YAGNI |
|
||||||
|
| Scrapers custom par plateforme | Maintenance impossible |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Roadmap
|
||||||
|
|
||||||
|
### Phase 1 — Base (recommandée)
|
||||||
|
|
||||||
|
- Route `/api/transcript/...`
|
||||||
|
- Module `transcript.mjs`
|
||||||
|
- UI #1 : affichage + sélecteur de langue
|
||||||
|
- Cache + rate-limit
|
||||||
|
- Tests
|
||||||
|
|
||||||
|
### Phase 2 — Confort
|
||||||
|
|
||||||
|
- UI #2 : clic sur une ligne = seek
|
||||||
|
- Branchement par provider
|
||||||
|
- Meilleure gestion des timecodes
|
||||||
|
|
||||||
|
### Phase 3 — Scalabilité
|
||||||
|
|
||||||
|
- Cache Redis
|
||||||
|
- Proxies rotatifs
|
||||||
|
- Monitoring avancé
|
||||||
|
- Support de nouveaux formats
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Annexes
|
||||||
|
|
||||||
|
### 13.1 Format `json3`
|
||||||
|
|
||||||
|
Exemple simplifié :
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"events": [
|
||||||
|
{
|
||||||
|
"tStartMs": 0,
|
||||||
|
"dDurationMs": 2500,
|
||||||
|
"segs": [{ "utf8": "Bonjour" }]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tStartMs": 2500,
|
||||||
|
"dDurationMs": 3100,
|
||||||
|
"segs": [{ "utf8": "Bienvenue" }]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.2 Format VTT
|
||||||
|
|
||||||
|
Exemple simplifié :
|
||||||
|
|
||||||
|
```vtt
|
||||||
|
WEBVTT
|
||||||
|
|
||||||
|
00:00:00.000 --> 00:00:02.500
|
||||||
|
Bonjour
|
||||||
|
|
||||||
|
00:00:02.500 --> 00:00:05.600
|
||||||
|
Bienvenue
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.3 Exemple de réponse complète
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"lang": "fr",
|
||||||
|
"available": true,
|
||||||
|
"languages": ["fr", "en", "es"],
|
||||||
|
"lines": [
|
||||||
|
{ "t": 0, "dur": 2.5, "text": "Bonjour à tous" },
|
||||||
|
{ "t": 2.5, "dur": 3.1, "text": "Bienvenue dans cette vidéo" },
|
||||||
|
{ "t": 5.6, "dur": 4.2, "text": "Aujourd'hui, nous allons voir..." }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.4 Arborescence des fichiers modifiés
|
||||||
|
|
||||||
|
```text
|
||||||
|
server/
|
||||||
|
index.mjs # route + cache + rate-limit
|
||||||
|
transcript.mjs # fonctions pures
|
||||||
|
tests/
|
||||||
|
transcript.test.mjs # tests
|
||||||
|
watch.component.ts # logique UI
|
||||||
|
watch.component.html # panneau Transcript
|
||||||
|
package.json # script test:transcript
|
||||||
|
README.md # roadmap + version
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Fin du guide.**
|
||||||
Generated
+36
@@ -34,6 +34,7 @@
|
|||||||
"rxjs": "^7.8.2",
|
"rxjs": "^7.8.2",
|
||||||
"tailwindcss": "latest",
|
"tailwindcss": "latest",
|
||||||
"youtube-dl-exec": "^3.0.0",
|
"youtube-dl-exec": "^3.0.0",
|
||||||
|
"youtubei.js": "^18.1.0",
|
||||||
"zone.js": "~0.15.1"
|
"zone.js": "~0.15.1"
|
||||||
},
|
},
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
@@ -965,6 +966,12 @@
|
|||||||
"node": ">=6.9.0"
|
"node": ">=6.9.0"
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"node_modules/@bufbuild/protobuf": {
|
||||||
|
"version": "2.15.0",
|
||||||
|
"resolved": "https://registry.npmjs.org/@bufbuild/protobuf/-/protobuf-2.15.0.tgz",
|
||||||
|
"integrity": "sha512-DAheWUkVr/SJTWCc+lg9dhY0eN4SaWlf4+bG1KzHeXbnqt0AfB/NX0Z+VunGlM1ki1B4zVvye27MpKh/svySUA==",
|
||||||
|
"license": "(Apache-2.0 AND BSD-3-Clause)"
|
||||||
|
},
|
||||||
"node_modules/@cspotcode/source-map-support": {
|
"node_modules/@cspotcode/source-map-support": {
|
||||||
"version": "0.8.1",
|
"version": "0.8.1",
|
||||||
"resolved": "https://registry.npmjs.org/@cspotcode/source-map-support/-/source-map-support-0.8.1.tgz",
|
"resolved": "https://registry.npmjs.org/@cspotcode/source-map-support/-/source-map-support-0.8.1.tgz",
|
||||||
@@ -5672,6 +5679,12 @@
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"node_modules/fflate": {
|
||||||
|
"version": "0.8.3",
|
||||||
|
"resolved": "https://registry.npmjs.org/fflate/-/fflate-0.8.3.tgz",
|
||||||
|
"integrity": "sha512-tbZNuJrLwGUp3zshBtdy4W+ORxZuIh8a5ilyIEQDC5rY1f3U20JMry0Ll3WBzU58EZKsEuJFXhb5gwv8CsPvgA==",
|
||||||
|
"license": "MIT"
|
||||||
|
},
|
||||||
"node_modules/ffmpeg-static": {
|
"node_modules/ffmpeg-static": {
|
||||||
"version": "5.2.0",
|
"version": "5.2.0",
|
||||||
"resolved": "https://registry.npmjs.org/ffmpeg-static/-/ffmpeg-static-5.2.0.tgz",
|
"resolved": "https://registry.npmjs.org/ffmpeg-static/-/ffmpeg-static-5.2.0.tgz",
|
||||||
@@ -6943,6 +6956,15 @@
|
|||||||
"integrity": "sha512-abv/qOcuPfk3URPfDzmZU1LKmuw8kT+0nIHvKrKgFrwifol/doWcdA4ZqsWQ8ENrFKkd67Mfpo/LovbIUsbt3w==",
|
"integrity": "sha512-abv/qOcuPfk3URPfDzmZU1LKmuw8kT+0nIHvKrKgFrwifol/doWcdA4ZqsWQ8ENrFKkd67Mfpo/LovbIUsbt3w==",
|
||||||
"license": "MIT"
|
"license": "MIT"
|
||||||
},
|
},
|
||||||
|
"node_modules/meriyah": {
|
||||||
|
"version": "7.3.3",
|
||||||
|
"resolved": "https://registry.npmjs.org/meriyah/-/meriyah-7.3.3.tgz",
|
||||||
|
"integrity": "sha512-uE5cnoNj+UYhoMdZDuymCzr5TzEuE3ZF2C4kn6Z76rLhkwnvZ/+6dUZnM41T3fifyuhW0+n8vK7eq6DKJxUfPA==",
|
||||||
|
"license": "ISC",
|
||||||
|
"engines": {
|
||||||
|
"node": ">=20.0.0"
|
||||||
|
}
|
||||||
|
},
|
||||||
"node_modules/methods": {
|
"node_modules/methods": {
|
||||||
"version": "1.1.2",
|
"version": "1.1.2",
|
||||||
"resolved": "https://registry.npmjs.org/methods/-/methods-1.1.2.tgz",
|
"resolved": "https://registry.npmjs.org/methods/-/methods-1.1.2.tgz",
|
||||||
@@ -9992,6 +10014,20 @@
|
|||||||
"node": ">= 18"
|
"node": ">= 18"
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"node_modules/youtubei.js": {
|
||||||
|
"version": "18.1.0",
|
||||||
|
"resolved": "https://registry.npmjs.org/youtubei.js/-/youtubei.js-18.1.0.tgz",
|
||||||
|
"integrity": "sha512-4Goo3ZeO6/kE8Mq67eC8Z9K1rijeqRwyKxf/YM4lqf8bB4r/s8JC0V3YFx7+l/AJR3LnAsulTsiIj+AFe0cduA==",
|
||||||
|
"funding": [
|
||||||
|
"https://github.com/sponsors/LuanRT"
|
||||||
|
],
|
||||||
|
"license": "MIT",
|
||||||
|
"dependencies": {
|
||||||
|
"@bufbuild/protobuf": "^2.0.0",
|
||||||
|
"fflate": "^0.8.2",
|
||||||
|
"meriyah": "^7.3.1"
|
||||||
|
}
|
||||||
|
},
|
||||||
"node_modules/zod": {
|
"node_modules/zod": {
|
||||||
"version": "3.25.76",
|
"version": "3.25.76",
|
||||||
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
|
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
|
||||||
|
|||||||
+9
-3
@@ -15,7 +15,12 @@
|
|||||||
"test:telemetry": "node ./server/tests/telemetry.test.mjs",
|
"test:telemetry": "node ./server/tests/telemetry.test.mjs",
|
||||||
"test:downloads": "node ./server/tests/download_jobs.test.mjs",
|
"test:downloads": "node ./server/tests/download_jobs.test.mjs",
|
||||||
"test:subscriptions": "node --loader ts-node/esm --experimental-specifier-resolution=node src/services/subscriptions.service.spec.ts",
|
"test:subscriptions": "node --loader ts-node/esm --experimental-specifier-resolution=node src/services/subscriptions.service.spec.ts",
|
||||||
"test:search": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search.service.spec.ts && node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search-components.spec.ts"
|
"test:search": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search.service.spec.ts && node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/search-components.spec.ts",
|
||||||
|
"test:suggest": "node --loader ts-node/esm --experimental-specifier-resolution=node src/app/search/suggest.spec.ts && node ./server/tests/suggest.test.mjs",
|
||||||
|
"test:transcript": "node --test server/tests/transcript.test.mjs",
|
||||||
|
"test:ytscrape": "node server/tests/youtube-scrape.test.mjs",
|
||||||
|
"test:ytinnertube": "node server/tests/youtube-innertube.test.mjs",
|
||||||
|
"ytdlp:update": "yt-dlp -U || python3 -m yt_dlp -U || echo \"yt-dlp update: installez yt-dlp puis relancez\""
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@angular/build": "^20.1.0",
|
"@angular/build": "^20.1.0",
|
||||||
@@ -44,12 +49,13 @@
|
|||||||
"rxjs": "^7.8.2",
|
"rxjs": "^7.8.2",
|
||||||
"tailwindcss": "latest",
|
"tailwindcss": "latest",
|
||||||
"youtube-dl-exec": "^3.0.0",
|
"youtube-dl-exec": "^3.0.0",
|
||||||
|
"youtubei.js": "18.1.0",
|
||||||
"zone.js": "~0.15.1"
|
"zone.js": "~0.15.1"
|
||||||
},
|
},
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
"@types/node": "^22.14.0",
|
"@types/node": "^22.14.0",
|
||||||
|
"ts-node": "^10.9.2",
|
||||||
"typescript": "~5.8.2",
|
"typescript": "~5.8.2",
|
||||||
"vite": "^6.2.0",
|
"vite": "^6.2.0"
|
||||||
"ts-node": "^10.9.2"
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -998,3 +998,74 @@ function subscriptionRowToDto(row) {
|
|||||||
channel,
|
channel,
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// -------------------- Step 17 : cache persistant recherche YouTube --------------------
|
||||||
|
function ensureYoutubeCacheTables() {
|
||||||
|
try {
|
||||||
|
db.exec(`CREATE TABLE IF NOT EXISTS youtube_search_cache (
|
||||||
|
q_hash TEXT PRIMARY KEY, q TEXT NOT NULL, payload_json TEXT NOT NULL,
|
||||||
|
source TEXT NOT NULL DEFAULT 'scrape', created_at INTEGER NOT NULL, expires_at INTEGER NOT NULL
|
||||||
|
);`);
|
||||||
|
db.exec(`CREATE INDEX IF NOT EXISTS idx_yt_cache_exp ON youtube_search_cache(expires_at);`);
|
||||||
|
db.exec(`CREATE TABLE IF NOT EXISTS youtube_metrics (
|
||||||
|
day TEXT PRIMARY KEY, scrape_calls INTEGER NOT NULL DEFAULT 0,
|
||||||
|
api_calls INTEGER NOT NULL DEFAULT 0, quota_units INTEGER NOT NULL DEFAULT 0, updated_at TEXT NOT NULL
|
||||||
|
);`);
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
ensureYoutubeCacheTables();
|
||||||
|
|
||||||
|
export function getCachedYoutubeSearch(qHash) {
|
||||||
|
try {
|
||||||
|
ensureYoutubeCacheTables();
|
||||||
|
const row = db.prepare(`SELECT payload_json AS payload, source, expires_at AS exp FROM youtube_search_cache WHERE q_hash = ?`).get(qHash);
|
||||||
|
if (!row) return null;
|
||||||
|
if (Date.now() >= Number(row.exp || 0)) {
|
||||||
|
try { db.prepare(`DELETE FROM youtube_search_cache WHERE q_hash = ?`).run(qHash); } catch {}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
try { return { items: JSON.parse(String(row.payload || '[]')), source: row.source }; } catch { return null; }
|
||||||
|
} catch { return null; }
|
||||||
|
}
|
||||||
|
|
||||||
|
export function setCachedYoutubeSearch(qHash, q, items, source, ttlMs) {
|
||||||
|
try {
|
||||||
|
ensureYoutubeCacheTables();
|
||||||
|
const now = Date.now();
|
||||||
|
db.prepare(`INSERT INTO youtube_search_cache (q_hash, q, payload_json, source, created_at, expires_at)
|
||||||
|
VALUES (?, ?, ?, ?, ?, ?)
|
||||||
|
ON CONFLICT(q_hash) DO UPDATE SET q=excluded.q, payload_json=excluded.payload_json,
|
||||||
|
source=excluded.source, created_at=excluded.created_at, expires_at=excluded.expires_at`)
|
||||||
|
.run(qHash, String(q || '').slice(0, 300), JSON.stringify(items || []), source, now, now + ttlMs);
|
||||||
|
// Cap : garde 2000 entrées les plus fraîches
|
||||||
|
try { db.exec(`DELETE FROM youtube_search_cache WHERE q_hash NOT IN (SELECT q_hash FROM youtube_search_cache ORDER BY expires_at DESC LIMIT 2000)`); } catch {}
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function pruneYoutubeCache() {
|
||||||
|
try { db.prepare(`DELETE FROM youtube_search_cache WHERE expires_at <= ?`).run(Date.now()); } catch {}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function incYoutubeMetrics({ scrapeCalls = 0, apiCalls = 0, quotaUnits = 0 } = {}) {
|
||||||
|
try {
|
||||||
|
ensureYoutubeCacheTables();
|
||||||
|
const day = new Date().toISOString().slice(0, 10);
|
||||||
|
db.prepare(`INSERT INTO youtube_metrics (day, scrape_calls, api_calls, quota_units, updated_at)
|
||||||
|
VALUES (?, ?, ?, ?, ?)
|
||||||
|
ON CONFLICT(day) DO UPDATE SET scrape_calls = scrape_calls + ?, api_calls = api_calls + ?,
|
||||||
|
quota_units = quota_units + ?, updated_at = excluded.updated_at`)
|
||||||
|
.run(day, scrapeCalls, apiCalls, quotaUnits, new Date().toISOString(), scrapeCalls, apiCalls, quotaUnits);
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getYoutubeMetricsToday() {
|
||||||
|
try {
|
||||||
|
ensureYoutubeCacheTables();
|
||||||
|
const day = new Date().toISOString().slice(0, 10);
|
||||||
|
return db.prepare(`SELECT * FROM youtube_metrics WHERE day = ?`).get(day) || { day, scrape_calls: 0, api_calls: 0, quota_units: 0 };
|
||||||
|
} catch { return { scrape_calls: 0, api_calls: 0, quota_units: 0 }; }
|
||||||
|
}
|
||||||
|
|
||||||
|
export function countYoutubeCacheRows() {
|
||||||
|
try { return db.prepare(`SELECT COUNT(1) AS n FROM youtube_search_cache`).get()?.n || 0; } catch { return 0; }
|
||||||
|
}
|
||||||
|
|||||||
+458
-3
@@ -7,12 +7,16 @@ import bcrypt from 'bcryptjs';
|
|||||||
import jwt from 'jsonwebtoken';
|
import jwt from 'jsonwebtoken';
|
||||||
import fs from 'node:fs';
|
import fs from 'node:fs';
|
||||||
import path from 'node:path';
|
import path from 'node:path';
|
||||||
import youtubedl from 'youtube-dl-exec';
|
import youtubedlPkg, { create as createYtDlp } from 'youtube-dl-exec';
|
||||||
|
import { execFile as execFileCb } from 'node:child_process';
|
||||||
|
import { promisify } from 'node:util';
|
||||||
|
import { fileURLToPath as serverFileURLToPath } from 'node:url';
|
||||||
import ffmpegPath from 'ffmpeg-static';
|
import ffmpegPath from 'ffmpeg-static';
|
||||||
import * as cheerio from 'cheerio';
|
import * as cheerio from 'cheerio';
|
||||||
import axios from 'axios';
|
import axios from 'axios';
|
||||||
import rumbleRouter from './rumble.mjs';
|
import rumbleRouter from './rumble.mjs';
|
||||||
import { providerRegistry, validateProviders } from './providers/registry.mjs';
|
import { providerRegistry, validateProviders } from './providers/registry.mjs';
|
||||||
|
import { pickTrack, parseTrackText, parseVtt, orderedTracks, normalizeTranscriptProvider, transcriptTrackExt, looksLikeHtmlError } from './transcript.mjs';
|
||||||
import {
|
import {
|
||||||
getUserByUsername,
|
getUserByUsername,
|
||||||
getUserById,
|
getUserById,
|
||||||
@@ -73,9 +77,53 @@ import {
|
|||||||
} from './db.mjs';
|
} from './db.mjs';
|
||||||
import { getChannelAdapter, setTwitchTokenProvider } from './providers/channel-registry.mjs';
|
import { getChannelAdapter, setTwitchTokenProvider } from './providers/channel-registry.mjs';
|
||||||
import { fetchChannelContent } from './providers/channel-content.mjs';
|
import { fetchChannelContent } from './providers/channel-content.mjs';
|
||||||
|
import { getSearchMode, getYtDlpBin, hasCookiesFile, metricsSnapshot } from './providers/youtube-common.mjs';
|
||||||
|
import { ytScrapeCacheStats } from './providers/youtube.mjs';
|
||||||
|
|
||||||
const app = express();
|
const app = express();
|
||||||
const PORT = Number(process.env.PORT || 4000);
|
const PORT = Number(process.env.PORT || 4000);
|
||||||
|
// yt-dlp: prefer the newest binary available. The copy bundled with
|
||||||
|
// youtube-dl-exec goes stale (YouTube then answers "The page needs to be
|
||||||
|
// reloaded" to dump-single-json), while a system install is usually fresher.
|
||||||
|
// `YT_DLP_PATH` wins, then `yt-dlp` on PATH, then the bundled binary.
|
||||||
|
const execFileAsync = promisify(execFileCb);
|
||||||
|
async function binaryVersion(bin) {
|
||||||
|
try {
|
||||||
|
const { stdout } = await execFileAsync(bin, ['--version'], { timeout: 15000 });
|
||||||
|
return String(stdout || '').trim().split('\n')[0].trim();
|
||||||
|
} catch {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let youtubedl = youtubedlPkg;
|
||||||
|
let ytDlpInfo = 'bundled';
|
||||||
|
try {
|
||||||
|
let bundledBin = null;
|
||||||
|
try {
|
||||||
|
const u = new URL('../node_modules/youtube-dl-exec/bin/', import.meta.url);
|
||||||
|
const cand = path.join(String(serverFileURLToPath(u)), process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp');
|
||||||
|
if (fs.existsSync(cand)) bundledBin = cand;
|
||||||
|
} catch {}
|
||||||
|
const bundledVer = bundledBin ? await binaryVersion(bundledBin) : null;
|
||||||
|
const candidates = [];
|
||||||
|
if (process.env.YT_DLP_PATH) candidates.push(process.env.YT_DLP_PATH);
|
||||||
|
candidates.push(process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp', 'yt-dlp');
|
||||||
|
let best = null;
|
||||||
|
for (const cand of candidates) {
|
||||||
|
if (!cand) continue;
|
||||||
|
const ver = await binaryVersion(cand);
|
||||||
|
if (ver && (!best || ver > best.ver)) best = { bin: cand, ver };
|
||||||
|
}
|
||||||
|
if (best && (!bundledVer || best.ver >= bundledVer)) {
|
||||||
|
youtubedl = createYtDlp(best.bin);
|
||||||
|
ytDlpInfo = `${best.bin} (${best.ver})`;
|
||||||
|
} else if (bundledVer) {
|
||||||
|
ytDlpInfo = `bundled (${bundledVer})`;
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
console.warn('[config] yt-dlp binary selection failed, using bundled:', e?.message || e);
|
||||||
|
}
|
||||||
|
console.log(`[config] yt-dlp: ${ytDlpInfo}`);
|
||||||
const IS_PROD = String(process.env.NODE_ENV || '').toLowerCase() === 'production';
|
const IS_PROD = String(process.env.NODE_ENV || '').toLowerCase() === 'production';
|
||||||
const JWT_SECRET = process.env.JWT_SECRET || 'dev-secret-change-me';
|
const JWT_SECRET = process.env.JWT_SECRET || 'dev-secret-change-me';
|
||||||
if (!process.env.JWT_SECRET) {
|
if (!process.env.JWT_SECRET) {
|
||||||
@@ -1006,6 +1054,14 @@ r.get('/peertube/:instance/*', async (req, res) => {
|
|||||||
});
|
});
|
||||||
|
|
||||||
// -------------------- Generic video details (GET) --------------------
|
// -------------------- Generic video details (GET) --------------------
|
||||||
|
// Cache mémoire pour les vidéos connexes InnerTube (watch-next) : TTL 1h, LRU 200.
|
||||||
|
const YT_RELATED_TTL_MS = Number(process.env.YT_RELATED_TTL_MS || 60 * 60 * 1000);
|
||||||
|
const ytRelatedCache = new Map();
|
||||||
|
function ytRelatedCacheSet(key, items) {
|
||||||
|
if (ytRelatedCache.has(key)) ytRelatedCache.delete(key);
|
||||||
|
ytRelatedCache.set(key, { ts: Date.now(), items });
|
||||||
|
while (ytRelatedCache.size > 200) { const o = ytRelatedCache.keys().next().value; if (o === undefined) break; ytRelatedCache.delete(o); }
|
||||||
|
}
|
||||||
// Returns metadata such as title, description, uploader, thumbnail, duration and views for a provider/videoId
|
// Returns metadata such as title, description, uploader, thumbnail, duration and views for a provider/videoId
|
||||||
// Supports query params similar to download endpoints: instance (PeerTube), slug (Odysee), sourceUrl (direct)
|
// Supports query params similar to download endpoints: instance (PeerTube), slug (Odysee), sourceUrl (direct)
|
||||||
r.get('/details/:provider/:videoId', async (req, res) => {
|
r.get('/details/:provider/:videoId', async (req, res) => {
|
||||||
@@ -1071,6 +1127,22 @@ r.get('/details/:provider/:videoId', async (req, res) => {
|
|||||||
url,
|
url,
|
||||||
type: 'video',
|
type: 'video',
|
||||||
};
|
};
|
||||||
|
// Step 18 : vidéos connexes façon SmartTube (watch-next InnerTube, best-effort, 0 quota).
|
||||||
|
// N'implique que YouTube ; toute erreur -> `related: []`, la réponse reste 200.
|
||||||
|
if (String(provider) === 'youtube' && req.query.related !== '0') {
|
||||||
|
try {
|
||||||
|
const { getRelatedViaInnerTube } = await import('./providers/youtube-innertube.mjs');
|
||||||
|
const relKey = `related:${videoId}`;
|
||||||
|
const cached = ytRelatedCache.get(relKey);
|
||||||
|
if (cached && (Date.now() - cached.ts) < YT_RELATED_TTL_MS) {
|
||||||
|
out.related = cached.items;
|
||||||
|
} else {
|
||||||
|
const items = await getRelatedViaInnerTube(videoId, 24).catch(() => []);
|
||||||
|
out.related = items;
|
||||||
|
ytRelatedCacheSet(relKey, items);
|
||||||
|
}
|
||||||
|
} catch { out.related = []; }
|
||||||
|
}
|
||||||
return res.json(out);
|
return res.json(out);
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
return res.status(500).json({ error: 'details_failed', details: String(e?.message || e) });
|
return res.status(500).json({ error: 'details_failed', details: String(e?.message || e) });
|
||||||
@@ -1684,7 +1756,7 @@ r.post('/telemetry/events', authMiddleware, telemetryLimiter, (req, res) => {
|
|||||||
const { event, meta } = req.body || {};
|
const { event, meta } = req.body || {};
|
||||||
if (!event || typeof event !== 'string') return res.status(400).json({ error: 'event_required' });
|
if (!event || typeof event !== 'string') return res.status(400).json({ error: 'event_required' });
|
||||||
// Whitelist known event names to keep the table clean
|
// Whitelist known event names to keep the table clean
|
||||||
const allowed = new Set(['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open']);
|
const allowed = new Set(['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open', 'suggest_shown', 'suggest_used']);
|
||||||
if (!allowed.has(event)) return res.status(400).json({ error: 'unknown_event' });
|
if (!allowed.has(event)) return res.status(400).json({ error: 'unknown_event' });
|
||||||
const row = insertTelemetryEvent({ userId: req.user.id, event, meta: (meta && typeof meta === 'object') ? meta : null });
|
const row = insertTelemetryEvent({ userId: req.user.id, event, meta: (meta && typeof meta === 'object') ? meta : null });
|
||||||
return res.status(201).json(row || { ok: true });
|
return res.status(201).json(row || { ok: true });
|
||||||
@@ -1704,7 +1776,7 @@ r.get('/telemetry/events', authMiddleware, (req, res) => {
|
|||||||
r.get('/telemetry/summary', authMiddleware, (req, res) => {
|
r.get('/telemetry/summary', authMiddleware, (req, res) => {
|
||||||
const since = typeof req.query.since === 'string' ? req.query.since : undefined;
|
const since = typeof req.query.since === 'string' ? req.query.since : undefined;
|
||||||
const events = {};
|
const events = {};
|
||||||
for (const name of ['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open']) {
|
for (const name of ['search_submit', 'provider_picker_open', 'provider_apply', 'at_autocomplete_use', 'quick_menu_open', 'suggest_shown', 'suggest_used']) {
|
||||||
events[name] = countTelemetryEvents({ event: name, since });
|
events[name] = countTelemetryEvents({ event: name, since });
|
||||||
}
|
}
|
||||||
return res.json({ events });
|
return res.json({ events });
|
||||||
@@ -1944,6 +2016,53 @@ r.get('/img/odysee', async (req, res) => {
|
|||||||
app.use('/api', r);
|
app.use('/api', r);
|
||||||
// Health endpoint for container checks
|
// Health endpoint for container checks
|
||||||
app.get('/api/health', (_req, res) => res.json({ status: 'ok' }));
|
app.get('/api/health', (_req, res) => res.json({ status: 'ok' }));
|
||||||
|
// Step 17 : observabilité YouTube (mode, yt-dlp, cache, quota). Aucun secret exposé.
|
||||||
|
app.get(['/healthz', '/api/healthz'], async (_req, res) => {
|
||||||
|
try {
|
||||||
|
const [{ getYoutubeMetricsToday, countYoutubeCacheRows }, common] =
|
||||||
|
await Promise.all([import('./db.mjs'), import('./providers/youtube-common.mjs')]);
|
||||||
|
let ytdlpVersion = null;
|
||||||
|
let resolvedBin = null;
|
||||||
|
try {
|
||||||
|
resolvedBin = await common.resolveYtDlpBin();
|
||||||
|
const { stdout } = await execFileAsync(resolvedBin, ['--version'], { timeout: 10000 });
|
||||||
|
ytdlpVersion = String(stdout || '').trim().split('\n')[0].trim() || null;
|
||||||
|
} catch {}
|
||||||
|
const keys = common.getYouTubeKeys();
|
||||||
|
const banned = [];
|
||||||
|
try { for (const [k, until] of ytKeyBans || []) if (Date.now() < until) banned.push(`...${String(k).slice(-4)}`); } catch {}
|
||||||
|
res.json({
|
||||||
|
status: 'ok',
|
||||||
|
youtube: {
|
||||||
|
mode: getSearchMode(),
|
||||||
|
ytdlp: { bin: resolvedBin || getYtDlpBin(), version: ytdlpVersion, info: ytDlpInfo, binOk: Boolean(ytdlpVersion) },
|
||||||
|
antiban: {
|
||||||
|
cookiesFile: hasCookiesFile(),
|
||||||
|
poToken: Boolean(String(process.env.YT_PO_TOKEN || '').trim()),
|
||||||
|
egressProxy: Boolean(String(process.env.YT_EGRESS_PROXY || '').trim()),
|
||||||
|
},
|
||||||
|
cache: { ...ytScrapeCacheStats(), sqliteRows: countYoutubeCacheRows() },
|
||||||
|
metrics: { ...metricsSnapshot(), today: getYoutubeMetricsToday() },
|
||||||
|
keys: { count: keys.length, banned },
|
||||||
|
},
|
||||||
|
});
|
||||||
|
} catch (e) {
|
||||||
|
res.status(500).json({ status: 'error', error: String(e?.message || e) });
|
||||||
|
}
|
||||||
|
});
|
||||||
|
// Step 17 : trending YouTube sans clé (scrape) avec fallback [] propre.
|
||||||
|
app.get('/api/trending', async (req, res) => {
|
||||||
|
try {
|
||||||
|
const provider = String(req.query.provider || 'yt');
|
||||||
|
const limit = Math.min(50, Math.max(1, Number(req.query.limit || 24)));
|
||||||
|
if (provider !== 'yt') return res.status(400).json({ error: 'only yt supported in phase 1' });
|
||||||
|
const { getTrendingViaScrape } = await import('./providers/youtube-scrape.mjs');
|
||||||
|
const items = await getTrendingViaScrape(limit);
|
||||||
|
return res.json({ provider, items });
|
||||||
|
} catch (e) {
|
||||||
|
return res.json({ provider: 'yt', items: [], error: String(e?.code || e?.message || 'trending_failed') });
|
||||||
|
}
|
||||||
|
});
|
||||||
// Alias to support Angular dev proxy paths in both dev and production builds
|
// Alias to support Angular dev proxy paths in both dev and production builds
|
||||||
app.use('/proxy/api', r);
|
app.use('/proxy/api', r);
|
||||||
// Mount dedicated Rumble router (browse, search, video)
|
// Mount dedicated Rumble router (browse, search, video)
|
||||||
@@ -2170,6 +2289,342 @@ app.get('/api/search', async (req, res) => {
|
|||||||
}
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// -------------------- Query typeahead suggestions (Step 15) --------------------
|
||||||
|
// GET /api/search/suggest?q=…&providers=yt,dm&limit=10 -> { q, groups: { yt: string[], dm: string[] } }
|
||||||
|
// Fan-out over provider `suggest()` handlers; providers without one degrade to [] (never 500).
|
||||||
|
const SUGGEST_CACHE_TTL_MS = Number(process.env.SUGGEST_CACHE_TTL_MS || 5 * 60 * 1000);
|
||||||
|
const SUGGEST_CACHE_MAX_ENTRIES = 500;
|
||||||
|
/** @type {Map<string, { ts: number, data: any }>} */
|
||||||
|
const suggestCache = new Map();
|
||||||
|
function suggestCacheGet(key) {
|
||||||
|
const hit = suggestCache.get(key);
|
||||||
|
if (!hit) return null;
|
||||||
|
if ((Date.now() - hit.ts) >= SUGGEST_CACHE_TTL_MS) {
|
||||||
|
suggestCache.delete(key);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
// LRU refresh
|
||||||
|
suggestCache.delete(key);
|
||||||
|
suggestCache.set(key, hit);
|
||||||
|
return hit.data;
|
||||||
|
}
|
||||||
|
function suggestCacheSet(key, data) {
|
||||||
|
if (suggestCache.has(key)) suggestCache.delete(key);
|
||||||
|
suggestCache.set(key, { ts: Date.now(), data });
|
||||||
|
while (suggestCache.size > SUGGEST_CACHE_MAX_ENTRIES) {
|
||||||
|
const oldest = suggestCache.keys().next().value;
|
||||||
|
suggestCache.delete(oldest);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const suggestLimiter = rateLimit({
|
||||||
|
windowMs: 60 * 1000,
|
||||||
|
max: Number(process.env.SUGGEST_RATE_LIMIT || 60),
|
||||||
|
standardHeaders: true,
|
||||||
|
legacyHeaders: false,
|
||||||
|
});
|
||||||
|
app.get('/api/search/suggest', suggestLimiter, async (req, res) => {
|
||||||
|
try {
|
||||||
|
const rawQ = typeof req.query.q === 'string' ? req.query.q : '';
|
||||||
|
const q = rawQ.trim();
|
||||||
|
if (q.length < 2) {
|
||||||
|
return res.status(400).json({ error: 'q is required and must be at least 2 characters long' });
|
||||||
|
}
|
||||||
|
const limit = Math.min(20, Math.max(1, Number(req.query.limit || 10)));
|
||||||
|
const requested = typeof req.query.providers === 'string' ? String(req.query.providers) : '';
|
||||||
|
const validProviders = validateProviders(requested);
|
||||||
|
const cacheKey = `suggest:${validProviders.join(',')}:${q.toLowerCase()}:${limit}`;
|
||||||
|
const cached = suggestCacheGet(cacheKey);
|
||||||
|
if (cached) return res.json(cached);
|
||||||
|
const results = await Promise.allSettled(
|
||||||
|
validProviders.map((providerId) => {
|
||||||
|
const mod = providerRegistry[providerId];
|
||||||
|
if (!mod || typeof mod.suggest !== 'function') return Promise.resolve([]);
|
||||||
|
return Promise.resolve().then(() => mod.suggest(q, { limit }));
|
||||||
|
})
|
||||||
|
);
|
||||||
|
const groups = {};
|
||||||
|
results.forEach((result, index) => {
|
||||||
|
const providerId = validProviders[index];
|
||||||
|
if (result.status === 'fulfilled' && Array.isArray(result.value)) {
|
||||||
|
groups[providerId] = result.value
|
||||||
|
.map((s) => String(s ?? '').trim())
|
||||||
|
.filter(Boolean)
|
||||||
|
.slice(0, limit);
|
||||||
|
} else {
|
||||||
|
groups[providerId] = [];
|
||||||
|
}
|
||||||
|
});
|
||||||
|
const data = { q, groups };
|
||||||
|
suggestCacheSet(cacheKey, data);
|
||||||
|
return res.json(data);
|
||||||
|
} catch (e) {
|
||||||
|
return res.status(500).json({ error: 'suggest_failed', details: String(e?.message || e) });
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
// -------------------- Video transcripts (Step 16, Phase 1) --------------------
|
||||||
|
// GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=
|
||||||
|
// -> { lang, available, languages, lines: [{ t, dur, text }] }
|
||||||
|
// One endpoint / one parser / one UI whatever the provider (yt-dlp subtitles).
|
||||||
|
const TRANSCRIPT_CACHE_TTL_MS = Number(process.env.TRANSCRIPT_CACHE_TTL || 24 * 60 * 60 * 1000);
|
||||||
|
const TRANSCRIPT_CACHE_MAX_ENTRIES = 200;
|
||||||
|
/** @type {Map<string, { ts: number, data: any }>} */
|
||||||
|
const transcriptCache = new Map();
|
||||||
|
function transcriptCacheGet(key) {
|
||||||
|
const hit = transcriptCache.get(key);
|
||||||
|
if (!hit) return null;
|
||||||
|
if ((Date.now() - hit.ts) >= TRANSCRIPT_CACHE_TTL_MS) {
|
||||||
|
transcriptCache.delete(key);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
transcriptCache.delete(key);
|
||||||
|
transcriptCache.set(key, hit);
|
||||||
|
return hit.data;
|
||||||
|
}
|
||||||
|
function transcriptCacheSet(key, data) {
|
||||||
|
if (transcriptCache.has(key)) transcriptCache.delete(key);
|
||||||
|
transcriptCache.set(key, { ts: Date.now(), data });
|
||||||
|
while (transcriptCache.size > TRANSCRIPT_CACHE_MAX_ENTRIES) {
|
||||||
|
const oldest = transcriptCache.keys().next().value;
|
||||||
|
transcriptCache.delete(oldest);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const transcriptLimiter = rateLimit({
|
||||||
|
windowMs: 60 * 1000,
|
||||||
|
max: Number(process.env.TRANSCRIPT_RATE_LIMIT || 10),
|
||||||
|
standardHeaders: true,
|
||||||
|
legacyHeaders: false,
|
||||||
|
// Always answer JSON (default handler sends an HTML/text page, which the
|
||||||
|
// Angular HttpClient cannot parse as JSON and surfaces as a raw SyntaxError).
|
||||||
|
handler: (req, res) => {
|
||||||
|
return res.status(429).json({ available: false, error: 'rate_limited' });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
/** Providers known NOT to expose subtitle tracks via yt-dlp (no transcript possible). */
|
||||||
|
const TRANSCRIPT_UNSUPPORTED_PROVIDERS = new Set(['twitch', 'odysee', 'rumble']);
|
||||||
|
|
||||||
|
/** Download subtitles via yt-dlp (handles YouTube impersonation + 429-prone
|
||||||
|
* translated tracks). Tries `langs` in order, returns the first non-empty
|
||||||
|
* parsed lines with the language that worked, or null. Bounded by a timeout
|
||||||
|
* (yt-dlp subtitle downloads can hang on .part files); partial results on
|
||||||
|
* disk are still scanned when yt-dlp exits non-zero. */
|
||||||
|
async function transcriptViaYtDlp(url, langs) {
|
||||||
|
const osMod = await import('node:os');
|
||||||
|
const dir = await fs.promises.mkdtemp(path.join(osMod.tmpdir(), 'newtube-transcript-'));
|
||||||
|
const scanDir = async (wanted) => {
|
||||||
|
let files = [];
|
||||||
|
try {
|
||||||
|
files = (await fs.promises.readdir(dir)).filter((f) => /\.vtt$/i.test(f) && !/\.part$/i.test(f));
|
||||||
|
} catch { return null; }
|
||||||
|
// Ignore stale/empty files
|
||||||
|
const nonEmpty = [];
|
||||||
|
for (const f of files) {
|
||||||
|
try {
|
||||||
|
const st = await fs.promises.stat(path.join(dir, f));
|
||||||
|
if (st.size > 0) nonEmpty.push(f);
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
// Prefer requested languages first
|
||||||
|
nonEmpty.sort((a, b) => {
|
||||||
|
const la = a.toLowerCase(), lb = b.toLowerCase();
|
||||||
|
const ia = wanted.findIndex((w) => la.includes(`.${w.toLowerCase()}.`) || la.endsWith(`.${w.toLowerCase()}.vtt`));
|
||||||
|
const ib = wanted.findIndex((w) => lb.includes(`.${w.toLowerCase()}.`) || lb.endsWith(`.${w.toLowerCase()}.vtt`));
|
||||||
|
return (ia === -1 ? 99 : ia) - (ib === -1 ? 99 : ib);
|
||||||
|
});
|
||||||
|
for (const file of nonEmpty) {
|
||||||
|
try {
|
||||||
|
const text = await fs.promises.readFile(path.join(dir, file), 'utf8');
|
||||||
|
const lines = parseVtt(text);
|
||||||
|
if (lines.length) {
|
||||||
|
const m = /\.([a-z]{2,3}(?:-[a-z]{2,4})?)\.vtt$/i.exec(file);
|
||||||
|
return { lines, lang: m ? m[1] : null };
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
return null;
|
||||||
|
};
|
||||||
|
try {
|
||||||
|
const wanted = Array.from(new Set((langs || []).map((l) => String(l || '').split('-')[0]).filter(Boolean)));
|
||||||
|
if (!wanted.includes('en')) wanted.push('en');
|
||||||
|
const subLangs = wanted.slice(0, 4).join(',');
|
||||||
|
const child = youtubedl(url, {
|
||||||
|
writeSub: true,
|
||||||
|
writeAutoSub: true,
|
||||||
|
subLangs,
|
||||||
|
subFormat: 'vtt/best',
|
||||||
|
skipDownload: true,
|
||||||
|
noWarnings: true,
|
||||||
|
noCheckCertificates: true,
|
||||||
|
noPlaylist: true,
|
||||||
|
output: path.join(dir, '%(id)s'),
|
||||||
|
});
|
||||||
|
const timer = setTimeout(() => { try { child.kill('SIGKILL'); } catch {} }, 90000);
|
||||||
|
try {
|
||||||
|
await child;
|
||||||
|
} catch (e) {
|
||||||
|
// Non-zero exit (e.g. one language 429'd) — partial files may still exist.
|
||||||
|
console.warn('[transcript] yt-dlp subtitle download exited non-zero:', String(e?.message || e).slice(0, 200));
|
||||||
|
} finally {
|
||||||
|
clearTimeout(timer);
|
||||||
|
}
|
||||||
|
return await scanDir(wanted);
|
||||||
|
} catch (e) {
|
||||||
|
console.warn('[transcript] yt-dlp subtitle download failed:', e?.message || e);
|
||||||
|
try {
|
||||||
|
const wanted = Array.from(new Set((langs || []).map((l) => String(l || '').split('-')[0]).filter(Boolean)));
|
||||||
|
return await scanDir(wanted);
|
||||||
|
} catch { return null; }
|
||||||
|
} finally {
|
||||||
|
try { await fs.promises.rm(dir, { recursive: true, force: true }); } catch {}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Fetch one timedtext track URL and parse it (any format). Returns [] on any failure. */
|
||||||
|
async function fetchTimedTextLines(track) {
|
||||||
|
const resp = await fetch(String(track.url), {
|
||||||
|
headers: {
|
||||||
|
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36',
|
||||||
|
Accept: 'application/json, text/vtt, text/*;q=0.9, */*;q=0.8',
|
||||||
|
'Accept-Language': 'en-US,en;q=0.9,fr;q=0.8',
|
||||||
|
},
|
||||||
|
signal: AbortSignal.timeout(15000),
|
||||||
|
});
|
||||||
|
if (!resp.ok) throw new Error(`track_fetch_failed:${resp.status}`);
|
||||||
|
const text = await resp.text();
|
||||||
|
if (!text || looksLikeHtmlError(text)) throw new Error('track_fetch_failed:html_error_page');
|
||||||
|
return parseTrackText(text, transcriptTrackExt(track));
|
||||||
|
}
|
||||||
|
|
||||||
|
app.get('/api/transcript/:provider/:videoId', transcriptLimiter, async (req, res) => {
|
||||||
|
try {
|
||||||
|
const { provider, videoId } = req.params;
|
||||||
|
const lang = String(req.query.lang || 'fr').slice(0, 12) || 'fr';
|
||||||
|
const normalized = normalizeTranscriptProvider(provider);
|
||||||
|
if (!normalized || !videoId) {
|
||||||
|
return res.status(400).json({ available: false, error: 'invalid_provider_or_video' });
|
||||||
|
}
|
||||||
|
const cacheKey = `transcript:${normalized}:${videoId}:${lang.toLowerCase()}`;
|
||||||
|
const cached = transcriptCacheGet(cacheKey);
|
||||||
|
if (cached) return res.json(cached);
|
||||||
|
// Source de découverte des pistes : InnerTube d'abord pour YouTube
|
||||||
|
// (mêmes URLs timedtext, sans spawn yt-dlp), yt-dlp sinon/en secours.
|
||||||
|
// YT_TRANSCRIPT_SOURCE=innertube-first (défaut) | ytdlp-only | innertube-only
|
||||||
|
const transcriptSource = String(process.env.YT_TRANSCRIPT_SOURCE || 'innertube-first').trim().toLowerCase();
|
||||||
|
const transcriptUrl = providerUrlFrom(normalized, String(videoId), {
|
||||||
|
instance: req.query.instance || undefined,
|
||||||
|
slug: req.query.slug || undefined,
|
||||||
|
sourceUrl: req.query.sourceUrl || undefined,
|
||||||
|
});
|
||||||
|
let meta = null;
|
||||||
|
if (normalized === 'youtube' && transcriptSource !== 'ytdlp-only') {
|
||||||
|
try {
|
||||||
|
const { getCaptionTracksViaInnerTube } = await import('./providers/youtube-innertube.mjs');
|
||||||
|
const cap = await getCaptionTracksViaInnerTube(String(videoId));
|
||||||
|
if (cap.trackCount > 0) {
|
||||||
|
meta = { subtitles: cap.subtitles, automatic_captions: cap.automatic_captions };
|
||||||
|
console.log(`[transcript] source=innertube tracks=${cap.trackCount} langs=${cap.languages.join(',')}`);
|
||||||
|
} else {
|
||||||
|
// InnerTube fait foi (même backend que le lecteur) : 0 piste = pas
|
||||||
|
// de sous-titres, sans payer un dump yt-dlp complet.
|
||||||
|
console.log('[transcript] source=innertube tracks=0 -> no_subtitles');
|
||||||
|
const empty = { lang: null, available: false, languages: [], lines: [], reason: 'no_subtitles' };
|
||||||
|
transcriptCacheSet(cacheKey, empty);
|
||||||
|
return res.json(empty);
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
console.warn('[transcript] innertube discovery failed, fallback yt-dlp:', e?.code || String(e?.message || e).slice(0, 120));
|
||||||
|
if (transcriptSource === 'innertube-only') {
|
||||||
|
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (!meta) {
|
||||||
|
try {
|
||||||
|
const raw = await youtubedl(transcriptUrl, { dumpSingleJson: true, skipDownload: true, noWarnings: true, noCheckCertificates: true });
|
||||||
|
meta = (typeof raw === 'string') ? JSON.parse(raw || '{}') : (raw || {});
|
||||||
|
} catch (e) {
|
||||||
|
console.error('[transcript] yt-dlp failed:', e?.message || e);
|
||||||
|
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const { track, languages, lang: chosenLang } = pickTrack(meta, lang);
|
||||||
|
if (!track) {
|
||||||
|
// Definitive absence: distinguish "provider never exposes subtitles"
|
||||||
|
// (twitch/odysee/rumble) from "this video has none" — cacheable 200s.
|
||||||
|
const reason = TRANSCRIPT_UNSUPPORTED_PROVIDERS.has(normalized) ? 'provider_unsupported' : 'no_subtitles';
|
||||||
|
const empty = { lang: null, available: false, languages: languages || [], lines: [], reason };
|
||||||
|
transcriptCacheSet(cacheKey, empty);
|
||||||
|
return res.json(empty);
|
||||||
|
}
|
||||||
|
// Try candidate tracks in order: requested language, then original (`en`),
|
||||||
|
// then everything else. Translated tracks are often rate-limited while the
|
||||||
|
// original still works — never fail on the first track alone.
|
||||||
|
// Bounded: hammering dozens of timedtext URLs only worsens YouTube 429s.
|
||||||
|
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
||||||
|
let lines = [];
|
||||||
|
let workedLang = chosenLang;
|
||||||
|
let sawRateLimit = false;
|
||||||
|
for (const cand of orderedTracks(meta, lang).slice(0, 5)) {
|
||||||
|
let attempt = 0;
|
||||||
|
for (;;) {
|
||||||
|
try {
|
||||||
|
const parsed = await fetchTimedTextLines(cand.track);
|
||||||
|
if (parsed && parsed.length) {
|
||||||
|
lines = parsed;
|
||||||
|
workedLang = cand.lang || chosenLang;
|
||||||
|
}
|
||||||
|
break;
|
||||||
|
} catch (e) {
|
||||||
|
const msg = String(e?.message || e);
|
||||||
|
console.warn('[transcript] track fetch failed:', cand.lang, msg);
|
||||||
|
// One retry after a short pause on transient 429s.
|
||||||
|
if (msg.includes(':429') && attempt === 0) {
|
||||||
|
sawRateLimit = true;
|
||||||
|
attempt += 1;
|
||||||
|
await sleep(1500);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (msg.includes(':429')) sawRateLimit = true;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (lines.length) break;
|
||||||
|
}
|
||||||
|
if (!lines || lines.length === 0) {
|
||||||
|
// Direct timedtext fetches failed (429 / Sorry pages / impersonation):
|
||||||
|
// let yt-dlp download the subtitles instead (handles all of the above).
|
||||||
|
try {
|
||||||
|
const viaDlp = await transcriptViaYtDlp(transcriptUrl, [lang, chosenLang].filter(Boolean));
|
||||||
|
if (viaDlp && viaDlp.lines && viaDlp.lines.length) {
|
||||||
|
lines = viaDlp.lines;
|
||||||
|
if (viaDlp.lang) workedLang = viaDlp.lang;
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
if (!lines || lines.length === 0) {
|
||||||
|
// Subtitle tracks exist but their content could not be retrieved or
|
||||||
|
// parsed (YouTube 429 / Sorry pages / impersonation): the video HAS
|
||||||
|
// subtitles, so this is transient — 502 with languages + retryable flag,
|
||||||
|
// never cached as "no subtitles".
|
||||||
|
return res.status(502).json({
|
||||||
|
available: false,
|
||||||
|
languages: languages || [],
|
||||||
|
lines: [],
|
||||||
|
error: 'transcript_temporarily_unavailable',
|
||||||
|
retryable: true,
|
||||||
|
rateLimited: sawRateLimit || undefined,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
const result = { lang: workedLang, available: true, languages: languages || [], lines };
|
||||||
|
transcriptCacheSet(cacheKey, result);
|
||||||
|
return res.json(result);
|
||||||
|
} catch (e) {
|
||||||
|
console.error('[transcript] unexpected error:', e?.message || e);
|
||||||
|
return res.status(502).json({ available: false, languages: [], lines: [], error: 'transcript_temporarily_unavailable', retryable: true });
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
// -------------------- Static Frontend (Angular build) --------------------
|
// -------------------- Static Frontend (Angular build) --------------------
|
||||||
const distRoot = path.join(process.cwd(), 'dist');
|
const distRoot = path.join(process.cwd(), 'dist');
|
||||||
const distBrowser = path.join(distRoot, 'browser');
|
const distBrowser = path.join(distRoot, 'browser');
|
||||||
|
|||||||
@@ -63,6 +63,16 @@ async function resolveYouTubeChannelId(externalId) {
|
|||||||
const raw = String(externalId || '').trim();
|
const raw = String(externalId || '').trim();
|
||||||
if (/^UC[\w-]{20,}$/.test(raw)) return raw;
|
if (/^UC[\w-]{20,}$/.test(raw)) return raw;
|
||||||
const handle = raw.replace(/^@/, '');
|
const handle = raw.replace(/^@/, '');
|
||||||
|
// 0) scrape sans clé (Step 17) : ne consomme aucun quota
|
||||||
|
try {
|
||||||
|
const { getSearchMode } = await import('./youtube-common.mjs');
|
||||||
|
const mode = getSearchMode();
|
||||||
|
if (mode !== 'api-only') {
|
||||||
|
const { resolveChannelIdViaScrape } = await import('./youtube-scrape.mjs');
|
||||||
|
const id = await resolveChannelIdViaScrape(raw);
|
||||||
|
if (id && /^UC[\w-]{20,}$/.test(id)) return id;
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
// 1) channels?forHandle (fonctionne encore pour beaucoup de chaînes)
|
// 1) channels?forHandle (fonctionne encore pour beaucoup de chaînes)
|
||||||
try {
|
try {
|
||||||
const data = await ytGet('channels', { part: 'id', forHandle: handle });
|
const data = await ytGet('channels', { part: 'id', forHandle: handle });
|
||||||
@@ -131,9 +141,32 @@ function ytTokenFor(key, page) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
async function ytContent(externalId, { type, page, limit, sort, q }) {
|
async function ytContent(externalId, { type, page, limit, sort, q }) {
|
||||||
const channelId = await resolveYouTubeChannelId(externalId);
|
|
||||||
const perPage = Math.min(Math.max(1, Number(limit || 24)), 50);
|
const perPage = Math.min(Math.max(1, Number(limit || 24)), 50);
|
||||||
const pageNum = Math.max(1, Number(page || 1));
|
const pageNum = Math.max(1, Number(page || 1));
|
||||||
|
// Step 17 : scrape-first sans clé (0 quota). Fallback API si bot-check/timeout.
|
||||||
|
try {
|
||||||
|
const { getSearchMode } = await import('./youtube-common.mjs');
|
||||||
|
const mode = getSearchMode();
|
||||||
|
if (mode !== 'api-only') {
|
||||||
|
const { fetchChannelViaScrape } = await import('./youtube-scrape.mjs');
|
||||||
|
// externalId brut (handle ou UC...) : le scrape gère les deux formes
|
||||||
|
const scraped = await fetchChannelViaScrape(externalId, { type, page: pageNum, limit: perPage });
|
||||||
|
if (Array.isArray(scraped?.items) && scraped.items.length) {
|
||||||
|
if (mode === 'scrape-only') return { ...scraped, total: null };
|
||||||
|
// scrape-first : retour direct si non vide
|
||||||
|
return { ...scraped, total: null };
|
||||||
|
}
|
||||||
|
// vide -> on tente l'API (chaîne à faible volume ou tab non supporté en scrape)
|
||||||
|
if (mode === 'scrape-only') return { items: [], nextPage: null };
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
console.warn('[channel-content/yt] scrape failed, fallback api:', e?.code || e?.message || e);
|
||||||
|
try {
|
||||||
|
const { getSearchMode } = await import('./youtube-common.mjs');
|
||||||
|
if (getSearchMode() === 'scrape-only') return { items: [], nextPage: null };
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
const channelId = await resolveYouTubeChannelId(externalId);
|
||||||
if (type === 'playlists') {
|
if (type === 'playlists') {
|
||||||
const key = ['pl', channelId, perPage].join('|');
|
const key = ['pl', channelId, perPage].join('|');
|
||||||
const token = ytTokenFor(key, pageNum);
|
const token = ytTokenFor(key, pageNum);
|
||||||
|
|||||||
@@ -0,0 +1,174 @@
|
|||||||
|
// Helpers partagés YouTube : clés API, mode de recherche, anti-ban yt-dlp, métriques.
|
||||||
|
// Extrait de youtube.mjs + index.mjs pour éviter la duplication (Step 17 P0).
|
||||||
|
import fs from 'node:fs';
|
||||||
|
import path from 'node:path';
|
||||||
|
import crypto from 'node:crypto';
|
||||||
|
import { execFile as execFileCb } from 'node:child_process';
|
||||||
|
import { promisify } from 'node:util';
|
||||||
|
|
||||||
|
export function getYouTubeKeys() {
|
||||||
|
const keys = [];
|
||||||
|
try {
|
||||||
|
const raw = process.env.YOUTUBE_API_KEYS;
|
||||||
|
if (raw && String(raw).trim() && !['undefined', 'null'].includes(String(raw).trim())) {
|
||||||
|
const s = String(raw).trim();
|
||||||
|
if (s.startsWith('[')) {
|
||||||
|
try {
|
||||||
|
const arr = JSON.parse(s);
|
||||||
|
if (Array.isArray(arr)) keys.push(...arr.map((v) => String(v || '').trim()).filter(Boolean));
|
||||||
|
} catch {}
|
||||||
|
} else {
|
||||||
|
keys.push(...s.split(',').map((v) => String(v || '').trim()).filter(Boolean));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
try {
|
||||||
|
const single = process.env.YOUTUBE_API_KEY;
|
||||||
|
if (single && String(single).trim()) keys.push(String(single).trim());
|
||||||
|
} catch {}
|
||||||
|
return Array.from(new Set(keys.filter(Boolean)));
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isKeyFailure(status, data) {
|
||||||
|
try {
|
||||||
|
const reason = data?.error?.errors?.[0]?.reason || '';
|
||||||
|
const message = String(data?.error?.message || '');
|
||||||
|
if (status === 400 && (reason === 'API_KEY_INVALID' || /api key (expired|invalid)/i.test(message))) return true;
|
||||||
|
if (status === 403 && /quota|rateLimit|dailyLimit|userRateLimit/i.test(`${reason} ${message}`)) return true;
|
||||||
|
} catch {}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Mode de recherche YT : innertube-first (défaut) | innertube-only | scrape-first | api-first | scrape-only | api-only */
|
||||||
|
export const YT_SEARCH_MODES = ['innertube-first', 'innertube-only', 'scrape-first', 'api-first', 'scrape-only', 'api-only'];
|
||||||
|
export function getSearchMode() {
|
||||||
|
const m = String(process.env.YT_SEARCH_MODE || 'innertube-first').trim().toLowerCase();
|
||||||
|
if (YT_SEARCH_MODES.includes(m)) return m;
|
||||||
|
return 'innertube-first';
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getScrapeTtlMs() {
|
||||||
|
const v = Number(process.env.YT_SCRAPE_TTL_MS || 30 * 60 * 1000);
|
||||||
|
return Number.isFinite(v) && v > 0 ? v : 30 * 60 * 1000;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getYtDlpTimeoutMs() {
|
||||||
|
const v = Number(process.env.YT_DLP_TIMEOUT_MS || 20000);
|
||||||
|
return Number.isFinite(v) && v >= 5000 ? v : 20000;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getYtDlpBin() {
|
||||||
|
const p = String(process.env.YT_DLP_PATH || '').trim();
|
||||||
|
if (p) return p;
|
||||||
|
return process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Résolution robuste du binaire yt-dlp (même ordre que server/index.mjs) :
|
||||||
|
* 1. YT_DLP_PATH (si le fichier existe)
|
||||||
|
* 2. `yt-dlp[.exe]` sur le PATH (probé via --version)
|
||||||
|
* 3. binaire bundled de `youtube-dl-exec` (node_modules/.../bin/, utilisé en Docker)
|
||||||
|
* Résultat mis en cache process-wide. Throw avec code `yt_scrape_no_binary` si introuvable.
|
||||||
|
*/
|
||||||
|
let _resolvedBin = null;
|
||||||
|
export function resetYtDlpBinCache() { _resolvedBin = null; }
|
||||||
|
|
||||||
|
function bundledYtDlpPath() {
|
||||||
|
try {
|
||||||
|
const exe = process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
|
||||||
|
// server/providers/ -> server/../node_modules/youtube-dl-exec/bin/
|
||||||
|
const cand = path.resolve(path.dirname(new URL(import.meta.url).pathname.replace(/^\/([A-Za-z]:)/, '$1')), '..', '..', 'node_modules', 'youtube-dl-exec', 'bin', exe);
|
||||||
|
if (fs.existsSync(cand)) return cand;
|
||||||
|
} catch {}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getBundledYtDlpPath() { return bundledYtDlpPath(); }
|
||||||
|
|
||||||
|
async function probeBin(bin) {
|
||||||
|
try {
|
||||||
|
const execFileAsync = promisify(execFileCb);
|
||||||
|
await execFileAsync(bin, ['--version'], { timeout: 10000 });
|
||||||
|
return true;
|
||||||
|
} catch (e) {
|
||||||
|
// ENOENT = binaire absent ; autre erreur (timeout...) = présent mais KO -> on le garde quand même
|
||||||
|
return e?.code !== 'ENOENT' && e?.errno !== 'ENOENT';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export async function resolveYtDlpBin() {
|
||||||
|
if (_resolvedBin) return _resolvedBin;
|
||||||
|
// 1) YT_DLP_PATH explicite
|
||||||
|
try {
|
||||||
|
const p = String(process.env.YT_DLP_PATH || '').trim();
|
||||||
|
if (p && fs.existsSync(p)) {
|
||||||
|
if (await probeBin(p)) { _resolvedBin = p; return p; }
|
||||||
|
} else if (p) {
|
||||||
|
console.warn(`[YT] YT_DLP_PATH introuvable : ${p} (fallback PATH/bundled)`);
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
// 2) PATH
|
||||||
|
const exe = process.platform === 'win32' ? 'yt-dlp.exe' : 'yt-dlp';
|
||||||
|
for (const cand of [exe, 'yt-dlp']) {
|
||||||
|
try {
|
||||||
|
if (await probeBin(cand)) { _resolvedBin = cand; return cand; }
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
// 3) bundled youtube-dl-exec
|
||||||
|
const bundled = bundledYtDlpPath();
|
||||||
|
if (bundled && await probeBin(bundled)) { _resolvedBin = bundled; return bundled; }
|
||||||
|
throw Object.assign(
|
||||||
|
new Error('yt-dlp introuvable (PATH, YT_DLP_PATH ni binaire bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
|
||||||
|
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Args anti-ban communs pour yt-dlp (cookies, PO-Token, proxy).
|
||||||
|
* Ne jamais logger les valeurs (cookies path exclu des logs verbeux).
|
||||||
|
*/
|
||||||
|
export function buildYtDlpExtraArgs() {
|
||||||
|
const args = [];
|
||||||
|
try {
|
||||||
|
const cookies = String(process.env.YT_COOKIES_FILE || '').trim();
|
||||||
|
if (cookies && fs.existsSync(cookies)) {
|
||||||
|
args.push('--cookies', cookies);
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
try {
|
||||||
|
const po = String(process.env.YT_PO_TOKEN || '').trim();
|
||||||
|
if (po) args.push('--extractor-args', `youtube:po_token=${po}`);
|
||||||
|
} catch {}
|
||||||
|
try {
|
||||||
|
const proxy = String(process.env.YT_EGRESS_PROXY || '').trim();
|
||||||
|
if (proxy) args.push('--proxy', proxy);
|
||||||
|
} catch {}
|
||||||
|
return args;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function hasCookiesFile() {
|
||||||
|
try {
|
||||||
|
const c = String(process.env.YT_COOKIES_FILE || '').trim();
|
||||||
|
return Boolean(c && fs.existsSync(c));
|
||||||
|
} catch { return false; }
|
||||||
|
}
|
||||||
|
|
||||||
|
export function hashSearchKey(parts) {
|
||||||
|
return crypto.createHash('sha256').update(String(parts)).digest('hex').slice(0, 32);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Métriques process-wide (remises à zéro au restart, persistées jour par jour en SQLite via db.mjs)
|
||||||
|
export const ytMetrics = {
|
||||||
|
scrapeCalls: 0,
|
||||||
|
scrapeHits: 0, // hit cache (mémoire ou SQLite)
|
||||||
|
apiCalls: 0,
|
||||||
|
quotaUnits: 0, // estimation search*100 + videos*1
|
||||||
|
fallbacks: 0,
|
||||||
|
scrapeErrors: 0,
|
||||||
|
innertubeCalls: 0,
|
||||||
|
innertubeErrors: 0,
|
||||||
|
};
|
||||||
|
|
||||||
|
export function metricsSnapshot() {
|
||||||
|
return { ...ytMetrics };
|
||||||
|
}
|
||||||
@@ -0,0 +1,331 @@
|
|||||||
|
// Couche InnerTube directe (façon SmartTube/MediaServiceCore) via youtubei.js.
|
||||||
|
// WEB client : search + continuations (pagination illimitée), watch-next (related),
|
||||||
|
// sans clé API ni quota. Le dispatcher bascule sur scrape/API si indisponible.
|
||||||
|
import { hashSearchKey, ytMetrics } from './youtube-common.mjs';
|
||||||
|
|
||||||
|
let sessionPromise = null;
|
||||||
|
let LogSilenced = false;
|
||||||
|
|
||||||
|
async function silenceLibNoise() {
|
||||||
|
if (LogSilenced) return;
|
||||||
|
LogSilenced = true;
|
||||||
|
try {
|
||||||
|
const { Log } = await import('youtubei.js');
|
||||||
|
// youtubei.js loggue en WARN chaque Text sans run assorti (bruit sur les titres) :
|
||||||
|
// on ne garde que les erreurs.
|
||||||
|
if (Log?.set_level && Log?.Level) {
|
||||||
|
const lvl = Log.Level.ERROR ?? Log.Level.WARNING ?? 3;
|
||||||
|
Log.set_level(lvl);
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Session InnerTube singleton (lazy). Throw code yt_innertube_unavailable si KO. */
|
||||||
|
export async function getSession() {
|
||||||
|
if (!sessionPromise) {
|
||||||
|
sessionPromise = (async () => {
|
||||||
|
try {
|
||||||
|
await silenceLibNoise();
|
||||||
|
const { Innertube } = await import('youtubei.js');
|
||||||
|
const gl = String(process.env.YT_INNERTUBE_GL || 'FR').trim() || 'FR';
|
||||||
|
const hl = String(process.env.YT_INNERTUBE_HL || 'fr').trim() || 'fr';
|
||||||
|
return await Innertube.create({ lang: hl, location: gl });
|
||||||
|
} catch (e) {
|
||||||
|
sessionPromise = null; // retry au prochain appel
|
||||||
|
throw Object.assign(
|
||||||
|
new Error(`InnerTube indisponible : ${String(e?.message || e).slice(0, 160)}`),
|
||||||
|
{ ytStatus: 502, code: 'yt_innertube_unavailable' },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
})();
|
||||||
|
}
|
||||||
|
return sessionPromise;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function resetSession() { sessionPromise = null; }
|
||||||
|
|
||||||
|
/** "81 973 vues" / "57K" / "1,2 M vues" -> number | undefined. */
|
||||||
|
export function parseViewsText(t) {
|
||||||
|
try {
|
||||||
|
const s = String(t || '').trim();
|
||||||
|
if (!s) return undefined;
|
||||||
|
const m = s.match(/([\d\s\u00a0.,]+)\s*([KMBkmb]|Mds|M|k|B)?/);
|
||||||
|
if (!m) return undefined;
|
||||||
|
const num = Number(m[1].replace(/[\s\u00a0]/g, '').replace(',', '.'));
|
||||||
|
if (!Number.isFinite(num)) return undefined;
|
||||||
|
const suffix = (m[2] || '').toLowerCase();
|
||||||
|
const mult = suffix === 'b' ? 1e9 : suffix === 'm' || suffix === 'mds' ? 1e6 : suffix === 'k' ? 1e3 : 1;
|
||||||
|
const v = Math.round(num * mult);
|
||||||
|
return v > 0 ? v : undefined;
|
||||||
|
} catch { return undefined; }
|
||||||
|
}
|
||||||
|
|
||||||
|
/** "2 minutes, 27 seconds" / "1 heure, 5 minutes" (label a11y) -> secondes. */
|
||||||
|
export function parseDurationLabel(label) {
|
||||||
|
try {
|
||||||
|
const s = String(label || '').toLowerCase();
|
||||||
|
if (!s) return undefined;
|
||||||
|
const get = (re) => { const m = s.match(re); return m ? Number(m[1]) : 0; };
|
||||||
|
const h = get(/(\d+)\s*(?:hours?|heures?)/);
|
||||||
|
const mnt = get(/(\d+)\s*(?:minutes?)/);
|
||||||
|
const sec = get(/(\d+)\s*(?:seconds?|secondes?)/);
|
||||||
|
const total = h * 3600 + mnt * 60 + sec;
|
||||||
|
return total > 0 ? total : undefined;
|
||||||
|
} catch { return undefined; }
|
||||||
|
}
|
||||||
|
|
||||||
|
function bestThumb(thumbs) {
|
||||||
|
try {
|
||||||
|
const arr = Array.isArray(thumbs) ? thumbs.filter((t) => t?.url) : [];
|
||||||
|
if (!arr.length) return undefined;
|
||||||
|
return arr[arr.length - 1].url;
|
||||||
|
} catch { return undefined; }
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Mappe un node InnerTube (Video, CompactVideo, GridVideo, LockupView, ReelItem,
|
||||||
|
* ShortsLockupView, PlaylistPanelVideo, WatchCardCompactVideo) vers Suggestion.
|
||||||
|
* Pur et testable offline. Les LockupView (nouveau renderer YouTube, utilisé
|
||||||
|
* notamment dans le watch-next façon SmartTube) ont une forme imbriquée propre.
|
||||||
|
*/
|
||||||
|
export function mapVideoNode(n) {
|
||||||
|
if (!n || typeof n !== 'object') return null;
|
||||||
|
try {
|
||||||
|
const type = String(n.type || '');
|
||||||
|
if (type === 'LockupView') return mapLockupView(n);
|
||||||
|
if (/playlist|channel|gridchannel|shelf|radio|show|album/i.test(type)
|
||||||
|
&& !/playlistpanelvideo|watchcard/i.test(type)) return null;
|
||||||
|
const id = n.id ? String(n.id) : null;
|
||||||
|
if (!id) return null;
|
||||||
|
const title = n.title?.text ?? (typeof n.title === 'string' ? n.title : '');
|
||||||
|
if (!title) return null;
|
||||||
|
const duration = n.duration?.seconds != null ? Number(n.duration.seconds) : undefined;
|
||||||
|
const author = n.author || n.uploader || null;
|
||||||
|
const authorName = author?.name ?? (typeof author === 'string' ? author : undefined);
|
||||||
|
const authorId = author?.id ? String(author.id) : undefined;
|
||||||
|
const views = parseViewsText(n.view_count?.text ?? n.view_count ?? n.views?.text);
|
||||||
|
const isShort = /short|reel/i.test(type) || (Number.isFinite(duration) && duration > 0 && duration <= 70 && /short/i.test(title) === false && /reel|short/i.test(type));
|
||||||
|
const badges = Array.isArray(n.badges) ? n.badges.map((b) => b?.label).filter(Boolean) : [];
|
||||||
|
return {
|
||||||
|
title: String(title),
|
||||||
|
id,
|
||||||
|
url: `https://www.youtube.com/watch?v=${id}`,
|
||||||
|
thumbnail: bestThumb(n.thumbnails),
|
||||||
|
uploaderName: authorName,
|
||||||
|
type: 'video',
|
||||||
|
...(Number.isFinite(duration) && duration > 0 ? { duration } : {}),
|
||||||
|
...(views !== undefined ? { views } : {}),
|
||||||
|
...(n.published?.text ? { publishedAt: String(n.published.text) } : {}),
|
||||||
|
...(authorId ? { channelId: authorId, channelExternalId: authorId, channelUrl: `https://www.youtube.com/channel/${authorId}` } : {}),
|
||||||
|
...(authorName ? { channelHandle: String(authorName) } : {}),
|
||||||
|
...(n.is_live ? { isLive: true } : {}),
|
||||||
|
...(isShort ? { isShort: true } : {}),
|
||||||
|
...(badges.length ? { badges } : {}),
|
||||||
|
};
|
||||||
|
} catch { return null; }
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Mappe un LockupView (renderer moderne : watch-next, search, shelves). */
|
||||||
|
export function mapLockupView(n) {
|
||||||
|
try {
|
||||||
|
const ct = String(n.content_type || '').toUpperCase();
|
||||||
|
if (ct && !['VIDEO', 'SHORT', 'MOVIE', 'LIVE'].includes(ct)) return null;
|
||||||
|
const id = n.content_id ? String(n.content_id)
|
||||||
|
: (n.renderer_context?.command_context?.on_tap?.payload?.videoId || null);
|
||||||
|
if (!id) return null;
|
||||||
|
const meta = n.metadata || {};
|
||||||
|
const title = meta.title?.text ?? (typeof meta.title === 'string' ? meta.title : '');
|
||||||
|
if (!title) return null;
|
||||||
|
// Vignette : content_image.image[] (prend la plus large)
|
||||||
|
let thumbnail;
|
||||||
|
try {
|
||||||
|
const imgs = (n.content_image?.image || []).filter((i) => i?.url);
|
||||||
|
thumbnail = imgs.sort((a, b) => (b.width || 0) - (a.width || 0))[0]?.url;
|
||||||
|
} catch {}
|
||||||
|
// Chaîne : 1ère ligne des metadata_rows, sinon label a11y "Go to channel X"
|
||||||
|
let uploaderName;
|
||||||
|
try {
|
||||||
|
const rows = meta.metadata?.metadata_rows || [];
|
||||||
|
const first = rows[0]?.metadata_parts?.[0]?.text;
|
||||||
|
uploaderName = first?.text ?? (typeof first === 'string' ? first : undefined);
|
||||||
|
} catch {}
|
||||||
|
if (!uploaderName) {
|
||||||
|
const a11y = String(meta.image?.a11y_label || '');
|
||||||
|
const m = a11y.match(/^(?:go to channel|aller sur la cha[îi]ne)\s+(.+)$/i);
|
||||||
|
if (m) uploaderName = m[1].trim();
|
||||||
|
}
|
||||||
|
// Vues : 2e ligne ("57K", "1,2 M vues"...), durée : label a11y
|
||||||
|
let views;
|
||||||
|
try {
|
||||||
|
const rows = meta.metadata?.metadata_rows || [];
|
||||||
|
const second = rows[1]?.metadata_parts?.[0]?.text;
|
||||||
|
views = parseViewsText(second?.text ?? second);
|
||||||
|
} catch {}
|
||||||
|
const duration = parseDurationLabel(n.renderer_context?.accessibility_context?.label);
|
||||||
|
return {
|
||||||
|
title: String(title),
|
||||||
|
id,
|
||||||
|
url: `https://www.youtube.com/watch?v=${id}`,
|
||||||
|
thumbnail,
|
||||||
|
uploaderName,
|
||||||
|
type: 'video',
|
||||||
|
...(duration !== undefined ? { duration } : {}),
|
||||||
|
...(views !== undefined ? { views } : {}),
|
||||||
|
...(uploaderName ? { channelHandle: String(uploaderName) } : {}),
|
||||||
|
...(ct === 'SHORT' ? { isShort: true } : {}),
|
||||||
|
...(ct === 'LIVE' ? { isLive: true } : {}),
|
||||||
|
};
|
||||||
|
} catch { return null; }
|
||||||
|
}
|
||||||
|
|
||||||
|
export function mapNodes(nodes) {
|
||||||
|
const seen = new Set();
|
||||||
|
const out = [];
|
||||||
|
for (const n of nodes || []) {
|
||||||
|
const m = mapVideoNode(n);
|
||||||
|
if (!m || seen.has(m.id)) continue;
|
||||||
|
seen.add(m.id);
|
||||||
|
out.push(m);
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Chaîne de continuations en mémoire : q_hash -> { feed, pages }.
|
||||||
|
// (SQLite persiste les pages déjà servies ; la chaîne évite de rejouer les pages 1..N-1.)
|
||||||
|
const chainCache = new Map();
|
||||||
|
const CHAIN_MAX = 50;
|
||||||
|
function chainGet(k) { const h = chainCache.get(k); if (h) { chainCache.delete(k); chainCache.set(k, h); } return h || null; }
|
||||||
|
function chainSet(k, v) {
|
||||||
|
if (chainCache.has(k)) chainCache.delete(k);
|
||||||
|
chainCache.set(k, v);
|
||||||
|
while (chainCache.size > CHAIN_MAX) { const o = chainCache.keys().next().value; if (o === undefined) break; chainCache.delete(o); }
|
||||||
|
}
|
||||||
|
export function innertubeChainStats() { return { chains: chainCache.size, max: CHAIN_MAX }; }
|
||||||
|
|
||||||
|
function sortParam(sort) {
|
||||||
|
const s = String(sort || 'relevance').toLowerCase();
|
||||||
|
if (s === 'date') return 'upload_date';
|
||||||
|
if (s === 'views') return 'view_count';
|
||||||
|
return 'relevance';
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Recherche InnerTube avec vraie pagination (continuations).
|
||||||
|
* Page 1 ~35-40 vidéos ; pages suivantes via getContinuation() fusionnées (façon SmartTube).
|
||||||
|
*/
|
||||||
|
export async function searchViaInnerTube(q, opts = {}) {
|
||||||
|
const query = String(q || '').trim();
|
||||||
|
if (query.length < 2) return [];
|
||||||
|
const limit = Math.min(50, Math.max(1, Number(opts?.limit || 24)));
|
||||||
|
const page = Math.min(10, Math.max(1, Number(opts?.page || 1)));
|
||||||
|
const sort = String(opts?.sort || 'relevance');
|
||||||
|
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
|
||||||
|
const yt = await getSession();
|
||||||
|
const key = `it|${hashSearchKey(`${query.toLowerCase()}|${sort}`)}`;
|
||||||
|
let entry = chainGet(key);
|
||||||
|
let feed = entry?.feed || null;
|
||||||
|
let collected = entry?.items ? [...entry.items] : [];
|
||||||
|
let pagesDone = entry?.pages || 0;
|
||||||
|
// Note : une "page" InnerTube fait ~17-20 vidéos quel que soit `limit`.
|
||||||
|
// On charge donc des continuations jusqu'à couvrir la fenêtre demandée
|
||||||
|
// [start, start+limit[ (façon SmartTube qui remplit son écran au fil des
|
||||||
|
// continuations), au lieu d'aligner 1 page API = 1 page UI.
|
||||||
|
const need = page * limit;
|
||||||
|
try {
|
||||||
|
if (!feed) {
|
||||||
|
feed = await yt.search(query, { sort_by: sortParam(sort) });
|
||||||
|
collected = mapNodes(feed.videos);
|
||||||
|
pagesDone = 1;
|
||||||
|
}
|
||||||
|
let guard = 0;
|
||||||
|
while (collected.length < need && feed?.has_continuation && guard < 12) {
|
||||||
|
guard++;
|
||||||
|
feed = await feed.getContinuation();
|
||||||
|
const fresh = mapNodes(feed.videos).filter((m) => !collected.some((c) => c.id === m.id));
|
||||||
|
collected = collected.concat(fresh);
|
||||||
|
pagesDone++;
|
||||||
|
if (!fresh.length) break;
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
if (e?.code === 'yt_innertube_unavailable') throw e;
|
||||||
|
throw Object.assign(new Error(`InnerTube search failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
|
||||||
|
}
|
||||||
|
if (feed) chainSet(key, { feed, items: collected, pages: pagesDone });
|
||||||
|
const start = (page - 1) * limit;
|
||||||
|
return collected.slice(start, start + limit);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Vidéos connexes façon SmartTube (endpoint watch-next) : ce qui alimente
|
||||||
|
* la colonne "À suivre" sous le lecteur.
|
||||||
|
*/
|
||||||
|
export async function getRelatedViaInnerTube(videoId, limit = 24) {
|
||||||
|
const id = String(videoId || '').trim();
|
||||||
|
if (!id) return [];
|
||||||
|
const n = Math.min(50, Math.max(1, Number(limit || 24)));
|
||||||
|
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
|
||||||
|
const yt = await getSession();
|
||||||
|
try {
|
||||||
|
const info = await yt.getInfo(id);
|
||||||
|
const feed = info?.watch_next_feed;
|
||||||
|
const arr = Array.isArray(feed) ? feed : (feed ? Array.from(feed) : []);
|
||||||
|
return mapNodes(arr).filter((m) => m.id !== id).slice(0, n);
|
||||||
|
} catch (e) {
|
||||||
|
if (e?.code === 'yt_innertube_unavailable') throw e;
|
||||||
|
throw Object.assign(new Error(`InnerTube related failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// -------------------- Transcripts via InnerTube --------------------
|
||||||
|
// `getInfo().captions.caption_tracks[]` expose les mêmes URLs timedtext que
|
||||||
|
// yt-dlp découvre via un dump complet (`subtitles`/`automatic_captions`).
|
||||||
|
// On adapte leur forme vers le format yt-dlp pour réutiliser pickTrack(),
|
||||||
|
// orderedTracks() et parseTrackText() de transcript.mjs sans les toucher.
|
||||||
|
|
||||||
|
function normCaptionLang(code) {
|
||||||
|
return String(code || '').trim().toLowerCase().replace(/_/g, '-');
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Adapte des caption tracks InnerTube brutes vers un pseudo dump yt-dlp
|
||||||
|
* `{ subtitles, automatic_captions }` (kind 'asr' = auto-généré).
|
||||||
|
* Pur et testable offline.
|
||||||
|
*/
|
||||||
|
export function mapCaptionTracks(captionTracks) {
|
||||||
|
const subtitles = {};
|
||||||
|
const automatic_captions = {};
|
||||||
|
for (const t of captionTracks || []) {
|
||||||
|
const base = String(t?.base_url || '');
|
||||||
|
if (!base) continue;
|
||||||
|
const lang = normCaptionLang(t?.language_code) || 'und';
|
||||||
|
const name = t?.name?.text ?? (typeof t?.name === 'string' ? t.name : lang);
|
||||||
|
const sep = base.includes('?') ? '&' : '?';
|
||||||
|
const track = { url: `${base}${sep}fmt=json3`, ext: 'json3', name: String(name) };
|
||||||
|
const dict = t?.kind === 'asr' ? automatic_captions : subtitles;
|
||||||
|
if (!Array.isArray(dict[lang])) dict[lang] = [];
|
||||||
|
if (!dict[lang].some((x) => x.url === track.url)) dict[lang].push(track);
|
||||||
|
}
|
||||||
|
const languages = Array.from(new Set([...Object.keys(subtitles), ...Object.keys(automatic_captions)]));
|
||||||
|
const trackCount = Object.values(subtitles).concat(Object.values(automatic_captions))
|
||||||
|
.reduce((n, arr) => n + arr.length, 0);
|
||||||
|
return { subtitles, automatic_captions, languages, trackCount };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Découvre les pistes de sous-titres d'une vidéo via InnerTube (0 quota,
|
||||||
|
* pas de spawn yt-dlp). Retourne un pseudo dump yt-dlp prêt pour pickTrack().
|
||||||
|
*/
|
||||||
|
export async function getCaptionTracksViaInnerTube(videoId) {
|
||||||
|
const id = String(videoId || '').trim();
|
||||||
|
if (!id) return { subtitles: {}, automatic_captions: {}, languages: [], trackCount: 0 };
|
||||||
|
ytMetrics.innertubeCalls = (ytMetrics.innertubeCalls || 0) + 1;
|
||||||
|
const yt = await getSession();
|
||||||
|
try {
|
||||||
|
const info = await yt.getInfo(id);
|
||||||
|
const raw = info?.captions?.caption_tracks || [];
|
||||||
|
return mapCaptionTracks(Array.isArray(raw) ? raw : Array.from(raw));
|
||||||
|
} catch (e) {
|
||||||
|
if (e?.code === 'yt_innertube_unavailable') throw e;
|
||||||
|
throw Object.assign(new Error(`InnerTube captions failed: ${String(e?.message || e).slice(0, 160)}`), { ytStatus: 502, code: 'yt_innertube_failed' });
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,197 @@
|
|||||||
|
// Recherche / channel / trending YouTube SANS clé API via yt-dlp (Step 17 P1).
|
||||||
|
// Binaire résolu via YT_DLP_PATH > PATH. Jamais de secret en log.
|
||||||
|
import { execFile as execFileCb } from 'node:child_process';
|
||||||
|
import { promisify } from 'node:util';
|
||||||
|
import { getYtDlpTimeoutMs, buildYtDlpExtraArgs, resolveYtDlpBin } from './youtube-common.mjs';
|
||||||
|
|
||||||
|
const execFileAsync = promisify(execFileCb);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Normalise une entrée yt-dlp flat-playlist vers Suggestion (même contrat que youtube.mjs).
|
||||||
|
* @param {any} e entrée yt-dlp
|
||||||
|
*/
|
||||||
|
export function mapFlatEntry(e) {
|
||||||
|
if (!e || typeof e !== 'object') return null;
|
||||||
|
const id = e.id || e.display_id || null;
|
||||||
|
if (!id) return null;
|
||||||
|
const thumbs = Array.isArray(e.thumbnails) ? e.thumbnails : [];
|
||||||
|
const thumb = thumbs.length
|
||||||
|
? (thumbs.find((t) => t?.url && (t.height || 0) >= 360)?.url || thumbs[thumbs.length - 1]?.url)
|
||||||
|
: (e.thumbnail || undefined);
|
||||||
|
const duration = e.duration != null ? Number(e.duration) : undefined;
|
||||||
|
const views = e.view_count != null ? Number(e.view_count) : (e.views != null ? Number(e.views) : undefined);
|
||||||
|
const channelId = e.channel_id || e.uploader_id || undefined;
|
||||||
|
const uploaderName = e.channel || e.uploader || undefined;
|
||||||
|
return {
|
||||||
|
title: e.title || '',
|
||||||
|
id: String(id),
|
||||||
|
url: String(id).startsWith('http') ? String(id) : `https://www.youtube.com/watch?v=${id}`,
|
||||||
|
thumbnail: thumb,
|
||||||
|
uploaderName,
|
||||||
|
type: 'video',
|
||||||
|
...(Number.isFinite(duration) && duration > 0 ? { duration } : {}),
|
||||||
|
...(Number.isFinite(views) ? { views } : {}),
|
||||||
|
...(e.timestamp ? { publishedAt: new Date(Number(e.timestamp) * 1000).toISOString() } : {}),
|
||||||
|
...(channelId ? { channelId: String(channelId), channelExternalId: String(channelId), channelUrl: `https://www.youtube.com/channel/${channelId}` } : {}),
|
||||||
|
...(uploaderName ? { channelHandle: String(uploaderName) } : {}),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Parse le JSON de `yt-dlp --dump-single-json --flat-playlist` (objet ou NDJSON fog). */
|
||||||
|
export function parseFlatPlaylistJson(raw) {
|
||||||
|
try {
|
||||||
|
const text = String(raw || '').trim();
|
||||||
|
if (!text) return [];
|
||||||
|
// Cas standard : un seul objet JSON avec .entries
|
||||||
|
try {
|
||||||
|
const obj = JSON.parse(text);
|
||||||
|
const entries = Array.isArray(obj?.entries) ? obj.entries : (Array.isArray(obj) ? obj : []);
|
||||||
|
return entries.map(mapFlatEntry).filter(Boolean);
|
||||||
|
} catch {
|
||||||
|
// Fallback NDJSON (une ligne = un JSON)
|
||||||
|
const out = [];
|
||||||
|
for (const line of text.split('\n')) {
|
||||||
|
const t = line.trim();
|
||||||
|
if (!t) continue;
|
||||||
|
try {
|
||||||
|
const o = JSON.parse(t);
|
||||||
|
const m = mapFlatEntry(o);
|
||||||
|
if (m) out.push(m);
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
} catch { return []; }
|
||||||
|
}
|
||||||
|
|
||||||
|
export function classifyScrapeError(e) {
|
||||||
|
if (e?.code === 'yt_scrape_no_binary') return e;
|
||||||
|
const msg = String(e?.message || e || '');
|
||||||
|
if (/ENOENT/i.test(msg) || e?.code === 'ENOENT' || e?.errno === 'ENOENT') {
|
||||||
|
return Object.assign(
|
||||||
|
new Error('yt-dlp introuvable (ni PATH ni bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
|
||||||
|
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (/bot|sign in to confirm|challenge|cookies|login/i.test(msg)) {
|
||||||
|
return Object.assign(new Error('YouTube bot-check (cookies/PO-Token requis)'), { ytStatus: 502, code: 'yt_scrape_bot_check' });
|
||||||
|
}
|
||||||
|
if (/timed out|timeout|ETIMEDOUT|killed|SIGKILL/i.test(msg)) {
|
||||||
|
return Object.assign(new Error('YouTube scrape timeout'), { ytStatus: 504, code: 'yt_scrape_timeout' });
|
||||||
|
}
|
||||||
|
if (/unable to|unsupported|not available|private|deleted/i.test(msg)) {
|
||||||
|
return Object.assign(new Error(`YouTube scrape upstream: ${msg.slice(0, 160)}`), { ytStatus: 502, code: 'yt_scrape_upstream' });
|
||||||
|
}
|
||||||
|
return Object.assign(new Error(`YouTube scrape failed: ${msg.slice(0, 200)}`), { ytStatus: 502, code: 'yt_scrape_failed' });
|
||||||
|
}
|
||||||
|
|
||||||
|
async function runYtDlpFlat(queryOrUrl, { limit = 10, playlistStart = 1 } = {}) {
|
||||||
|
const bin = await resolveYtDlpBin();
|
||||||
|
const timeout = getYtDlpTimeoutMs();
|
||||||
|
const end = playlistStart + Math.max(1, Math.min(50, Number(limit || 10))) - 1;
|
||||||
|
const args = [
|
||||||
|
'--dump-single-json',
|
||||||
|
'--flat-playlist',
|
||||||
|
'--no-warnings',
|
||||||
|
'--no-check-certificates',
|
||||||
|
'--skip-download',
|
||||||
|
'--no-playlist',
|
||||||
|
'--playlist-start', String(playlistStart),
|
||||||
|
'--playlist-end', String(end),
|
||||||
|
...buildYtDlpExtraArgs(),
|
||||||
|
queryOrUrl,
|
||||||
|
];
|
||||||
|
try {
|
||||||
|
const { stdout } = await execFileAsync(bin, args, { timeout, maxBuffer: 16 * 1024 * 1024 });
|
||||||
|
return parseFlatPlaylistJson(stdout);
|
||||||
|
} catch (e) {
|
||||||
|
// yt-dlp peut écrire du JSON partiel sur stdout même en exit non-zero
|
||||||
|
const partial = e?.stdout ? parseFlatPlaylistJson(String(e.stdout)) : [];
|
||||||
|
if (partial.length) return partial;
|
||||||
|
throw classifyScrapeError(e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Recherche sans clé. sort: relevance|date|views */
|
||||||
|
export async function searchViaScrape(q, opts = {}) {
|
||||||
|
const query = String(q || '').trim();
|
||||||
|
if (query.length < 2) return [];
|
||||||
|
const limit = Math.min(50, Math.max(1, Number(opts?.limit || 10)));
|
||||||
|
const page = Math.max(1, Number(opts?.page || 1));
|
||||||
|
const sort = String(opts?.sort || 'relevance').toLowerCase();
|
||||||
|
const perPage = limit;
|
||||||
|
const start = (page - 1) * perPage + 1;
|
||||||
|
let prefix = 'ytsearch';
|
||||||
|
if (sort === 'date') prefix = 'ytsearchdate';
|
||||||
|
// views : pas de préfixe natif stable -> ytsearch + tri local
|
||||||
|
const target = `${prefix}${perPage}:${query}`;
|
||||||
|
let items = await runYtDlpFlat(target, { limit: perPage, playlistStart: start });
|
||||||
|
if (sort === 'views') items = [...items].sort((a, b) => (b.views || 0) - (a.views || 0));
|
||||||
|
return items.slice(0, perPage);
|
||||||
|
}
|
||||||
|
|
||||||
|
const CHANNEL_TABS = { videos: 'videos', shorts: 'shorts', streams: 'streams', live: 'streams', playlists: 'playlists' };
|
||||||
|
|
||||||
|
/** Contenu chaîne sans clé. type: videos|shorts|playlists|live */
|
||||||
|
export async function fetchChannelViaScrape(externalId, { type = 'videos', page = 1, limit = 24 } = {}) {
|
||||||
|
const raw = String(externalId || '').trim().replace(/^@/, '');
|
||||||
|
if (!raw) return { items: [], nextPage: null };
|
||||||
|
const tab = CHANNEL_TABS[String(type)] || 'videos';
|
||||||
|
const perPage = Math.min(50, Math.max(1, Number(limit || 24)));
|
||||||
|
const pageNum = Math.max(1, Number(page || 1));
|
||||||
|
const start = (pageNum - 1) * perPage + 1;
|
||||||
|
// UC... -> /channel/UC.../tab, sinon -> /@/handle/tab
|
||||||
|
const base = /^UC[\w-]{20,}$/.test(raw)
|
||||||
|
? `https://www.youtube.com/channel/${raw}/${tab}`
|
||||||
|
: `https://www.youtube.com/@${raw}/${tab}`;
|
||||||
|
const items = await runYtDlpFlat(base, { limit: perPage, playlistStart: start });
|
||||||
|
// playlists renvoient des ids PL... : on les garde tels quels avec un shape playlist
|
||||||
|
if (tab === 'playlists') {
|
||||||
|
const mapped = items.map((s) => ({ id: s.id, title: s.title, thumbnail: s.thumbnail, videoCount: null, updatedAt: s.publishedAt || null }));
|
||||||
|
return { items: mapped, nextPage: items.length >= perPage ? pageNum + 1 : null };
|
||||||
|
}
|
||||||
|
let filtered = items;
|
||||||
|
if (type === 'shorts') filtered = items.filter((s) => !s.duration || s.duration <= 70);
|
||||||
|
return { items: filtered, nextPage: items.length >= perPage ? pageNum + 1 : null };
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Trending sans clé. Les onglets /feed/trending ont été retirés côté YouTube
|
||||||
|
* (redirect home -> erreur tab) ; on tente les tabs puis repli ytsearch trié par vues. */
|
||||||
|
export async function getTrendingViaScrape(limit = 24) {
|
||||||
|
const n = Math.min(50, Math.max(1, Number(limit || 24)));
|
||||||
|
const tabs = [
|
||||||
|
'https://www.youtube.com/feed/trending',
|
||||||
|
'https://www.youtube.com/trending',
|
||||||
|
];
|
||||||
|
for (const url of tabs) {
|
||||||
|
try {
|
||||||
|
const items = await runYtDlpFlat(url, { limit: n, playlistStart: 1 });
|
||||||
|
if (items.length) return items;
|
||||||
|
} catch {}
|
||||||
|
}
|
||||||
|
// Repli : recherche générique triée par vues (0 quota, toujours disponible)
|
||||||
|
return runYtDlpFlat(`ytsearch${n}:top trending videos world`, { limit: n, playlistStart: 1 });
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Résout un channel_id via scrape (imprime channel_id sans appel API). */
|
||||||
|
export async function resolveChannelIdViaScrape(externalId) {
|
||||||
|
const raw = String(externalId || '').trim();
|
||||||
|
if (/^UC[\w-]{20,}$/.test(raw)) return raw;
|
||||||
|
const handle = raw.replace(/^@/, '');
|
||||||
|
const bin = await resolveYtDlpBin().catch(() => null);
|
||||||
|
if (!bin) {
|
||||||
|
throw Object.assign(
|
||||||
|
new Error('yt-dlp introuvable (ni PATH ni bundled). Installez yt-dlp ou définissez YT_DLP_PATH.'),
|
||||||
|
{ ytStatus: 503, code: 'yt_scrape_no_binary' },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
const timeout = Math.min(15000, getYtDlpTimeoutMs());
|
||||||
|
try {
|
||||||
|
const { stdout } = await execFileAsync(bin,
|
||||||
|
['--print', '%(channel_id)s', '--no-warnings', '--skip-download', '--playlist-end', '1', ...buildYtDlpExtraArgs(), `https://www.youtube.com/@${handle}/videos`],
|
||||||
|
{ timeout });
|
||||||
|
const id = String(stdout || '').trim().split('\n')[0].trim();
|
||||||
|
if (/^UC[\w-]{20,}$/.test(id)) return id;
|
||||||
|
} catch {}
|
||||||
|
return raw;
|
||||||
|
}
|
||||||
+268
-143
@@ -9,6 +9,12 @@
|
|||||||
* @property {string=} uploaderName
|
* @property {string=} uploaderName
|
||||||
* @property {string=} type
|
* @property {string=} type
|
||||||
*/
|
*/
|
||||||
|
import {
|
||||||
|
getYouTubeKeys, isKeyFailure, getSearchMode, getScrapeTtlMs,
|
||||||
|
hashSearchKey, ytMetrics,
|
||||||
|
} from './youtube-common.mjs';
|
||||||
|
import { searchViaScrape } from './youtube-scrape.mjs';
|
||||||
|
import { searchViaInnerTube } from './youtube-innertube.mjs';
|
||||||
|
|
||||||
function parseISODurationToSeconds(iso) {
|
function parseISODurationToSeconds(iso) {
|
||||||
if (typeof iso !== 'string' || !iso) return 0;
|
if (typeof iso !== 'string' || !iso) return 0;
|
||||||
@@ -20,50 +26,6 @@ function parseISODurationToSeconds(iso) {
|
|||||||
return (hours * 3600) + (minutes * 60) + seconds;
|
return (hours * 3600) + (minutes * 60) + seconds;
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
|
||||||
* Resolve configured YouTube API keys.
|
|
||||||
* Accepts YOUTUBE_API_KEYS as JSON array ('["k1","k2"]') or CSV ('k1,k2'),
|
|
||||||
* plus the legacy single YOUTUBE_API_KEY as fallback. De-duplicated.
|
|
||||||
* @returns {string[]}
|
|
||||||
*/
|
|
||||||
function getYouTubeKeys() {
|
|
||||||
const keys = [];
|
|
||||||
try {
|
|
||||||
const raw = process.env.YOUTUBE_API_KEYS;
|
|
||||||
if (raw && String(raw).trim() && String(raw).trim() !== 'undefined' && String(raw).trim() !== 'null') {
|
|
||||||
const s = String(raw).trim();
|
|
||||||
if (s.startsWith('[')) {
|
|
||||||
try {
|
|
||||||
const arr = JSON.parse(s);
|
|
||||||
if (Array.isArray(arr)) keys.push(...arr.map(v => String(v || '').trim()).filter(Boolean));
|
|
||||||
} catch {}
|
|
||||||
} else {
|
|
||||||
keys.push(...s.split(',').map(v => String(v || '').trim()).filter(Boolean));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
} catch {}
|
|
||||||
try {
|
|
||||||
const single = process.env.YOUTUBE_API_KEY;
|
|
||||||
if (single && String(single).trim()) keys.push(String(single).trim());
|
|
||||||
} catch {}
|
|
||||||
return Array.from(new Set(keys.filter(Boolean)));
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Classify a YouTube API failure as retryable with another key.
|
|
||||||
* - 400 API_KEY_INVALID / "API key expired" -> key is dead, try next.
|
|
||||||
* - 403 quotaExceeded / rateLimitExceeded / dailyLimitExceeded -> quota, try next.
|
|
||||||
*/
|
|
||||||
function isKeyFailure(status, data) {
|
|
||||||
try {
|
|
||||||
const reason = data?.error?.errors?.[0]?.reason || '';
|
|
||||||
const message = String(data?.error?.message || '');
|
|
||||||
if (status === 400 && (reason === 'API_KEY_INVALID' || /api key (expired|invalid)/i.test(message))) return true;
|
|
||||||
if (status === 403 && /quota|rateLimit|dailyLimit|userRateLimit/i.test(reason + ' ' + message)) return true;
|
|
||||||
} catch {}
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
|
|
||||||
/** GET JSON from YouTube, rotating through keys on key failures. */
|
/** GET JSON from YouTube, rotating through keys on key failures. */
|
||||||
async function ytFetchJson(base, paramsWithoutKey) {
|
async function ytFetchJson(base, paramsWithoutKey) {
|
||||||
const keys = getYouTubeKeys();
|
const keys = getYouTubeKeys();
|
||||||
@@ -76,13 +38,19 @@ async function ytFetchJson(base, paramsWithoutKey) {
|
|||||||
for (const key of keys) {
|
for (const key of keys) {
|
||||||
const params = new URLSearchParams(paramsWithoutKey);
|
const params = new URLSearchParams(paramsWithoutKey);
|
||||||
params.set('key', key);
|
params.set('key', key);
|
||||||
|
ytMetrics.apiCalls++;
|
||||||
|
try {
|
||||||
|
const { incYoutubeMetrics } = await import('../db.mjs').catch(() => ({}));
|
||||||
|
if (String(base).includes('/search')) incYoutubeMetrics?.({ apiCalls: 1, quotaUnits: 100 });
|
||||||
|
else incYoutubeMetrics?.({ apiCalls: 1, quotaUnits: 1 });
|
||||||
|
ytMetrics.quotaUnits += String(base).includes('/search') ? 100 : 1;
|
||||||
|
} catch {}
|
||||||
const resp = await fetch(`${base}?${params.toString()}`);
|
const resp = await fetch(`${base}?${params.toString()}`);
|
||||||
const data = await resp.json().catch(() => ({}));
|
const data = await resp.json().catch(() => ({}));
|
||||||
if (resp.ok) return data;
|
if (resp.ok) return data;
|
||||||
lastStatus = resp.status;
|
lastStatus = resp.status;
|
||||||
lastData = data;
|
lastData = data;
|
||||||
lastError = new Error(`YouTube API error: ${resp.status} ${data?.error?.message || ''}`.trim());
|
lastError = new Error(`YouTube API error: ${resp.status} ${data?.error?.message || ''}`.trim());
|
||||||
// Only rotate to the next key on key-attributable failures; otherwise fail fast.
|
|
||||||
if (!isKeyFailure(resp.status, data)) break;
|
if (!isKeyFailure(resp.status, data)) break;
|
||||||
console.warn(`[YouTube] key ...${String(key).slice(-4)} failed (${resp.status}), trying next key`);
|
console.warn(`[YouTube] key ...${String(key).slice(-4)} failed (${resp.status}), trying next key`);
|
||||||
}
|
}
|
||||||
@@ -92,121 +60,278 @@ async function ytFetchJson(base, paramsWithoutKey) {
|
|||||||
throw err;
|
throw err;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse a `suggestqueries.google.com` (client=youtube) response body into plain strings.
|
||||||
|
* The endpoint answers either bare JSON (`["q",["s1","s2"],...]`) or wrapped in
|
||||||
|
* `window.google.ac.h(...)`. Never throws: unparsable bodies yield [].
|
||||||
|
* @param {string} text
|
||||||
|
* @returns {string[]}
|
||||||
|
*/
|
||||||
|
export function parseYoutubeSuggestResponse(text) {
|
||||||
|
try {
|
||||||
|
const raw = String(text || '').trim();
|
||||||
|
if (!raw) return [];
|
||||||
|
let payload = raw;
|
||||||
|
const start = raw.indexOf('(');
|
||||||
|
const end = raw.lastIndexOf(')');
|
||||||
|
if (start !== -1 && end > start) payload = raw.slice(start + 1, end);
|
||||||
|
const parsed = JSON.parse(payload);
|
||||||
|
const candidates = Array.isArray(parsed) && Array.isArray(parsed[1]) ? parsed[1] : [];
|
||||||
|
return candidates
|
||||||
|
.map((s) => {
|
||||||
|
if (Array.isArray(s)) s = s[0];
|
||||||
|
return (typeof s === 'string' ? s : String(s ?? '')).trim();
|
||||||
|
})
|
||||||
|
.filter(Boolean)
|
||||||
|
.slice(0, 20);
|
||||||
|
} catch {
|
||||||
|
return [];
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Mémoire LRU process-wide pour le dispatcher (complète SQLite)
|
||||||
|
const memCache = new Map();
|
||||||
|
const MEM_MAX = Number(process.env.YT_SCRAPE_MEM_MAX || 300);
|
||||||
|
function memGet(k, ttl) {
|
||||||
|
const hit = memCache.get(k);
|
||||||
|
if (!hit) return null;
|
||||||
|
if ((Date.now() - hit.ts) >= ttl) { memCache.delete(k); return null; }
|
||||||
|
memCache.delete(k); memCache.set(k, hit);
|
||||||
|
return hit.items;
|
||||||
|
}
|
||||||
|
function memSet(k, items) {
|
||||||
|
if (memCache.has(k)) memCache.delete(k);
|
||||||
|
memCache.set(k, { ts: Date.now(), items });
|
||||||
|
while (memCache.size > MEM_MAX) { const o = memCache.keys().next().value; if (o === undefined) break; memCache.delete(o); }
|
||||||
|
}
|
||||||
|
export function ytScrapeCacheStats() {
|
||||||
|
return { memEntries: memCache.size, memMax: MEM_MAX };
|
||||||
|
}
|
||||||
|
|
||||||
|
async function searchViaApi(q, { limit = 10, page = 1, sort = 'relevance' } = {}) {
|
||||||
|
const keys = getYouTubeKeys();
|
||||||
|
if (!keys.length) {
|
||||||
|
throw Object.assign(new Error('YOUTUBE_API_KEY not configured'), { ytStatus: 503, code: 'youtube_api_key_unavailable' });
|
||||||
|
}
|
||||||
|
let order = 'relevance';
|
||||||
|
if (sort === 'date') order = 'date';
|
||||||
|
else if (sort === 'views') order = 'viewCount';
|
||||||
|
const perPage = Math.min(Math.max(1, Number(limit || 10)), 50);
|
||||||
|
const targetPage = Math.max(1, Number(page || 1));
|
||||||
|
let pageToken = '';
|
||||||
|
let currentPage = 1;
|
||||||
|
let lastItems = [];
|
||||||
|
while (currentPage <= targetPage) {
|
||||||
|
const params = {
|
||||||
|
part: 'snippet', q, type: 'video', maxResults: String(perPage), order,
|
||||||
|
videoEmbeddable: 'true', safeSearch: 'moderate',
|
||||||
|
};
|
||||||
|
if (pageToken) params.pageToken = pageToken;
|
||||||
|
const data = await ytFetchJson('https://www.googleapis.com/youtube/v3/search', params);
|
||||||
|
if (currentPage === targetPage) { lastItems = Array.isArray(data.items) ? data.items : []; break; }
|
||||||
|
const next = data.nextPageToken;
|
||||||
|
if (!next) { lastItems = []; break; }
|
||||||
|
pageToken = String(next);
|
||||||
|
currentPage++;
|
||||||
|
}
|
||||||
|
const videoIds = (lastItems || []).map((item) => item?.id?.videoId).filter(Boolean);
|
||||||
|
const detailsMap = new Map();
|
||||||
|
if (videoIds.length > 0) {
|
||||||
|
try {
|
||||||
|
const detailsData = await ytFetchJson('https://www.googleapis.com/youtube/v3/videos', {
|
||||||
|
part: 'contentDetails,statistics,status', id: videoIds.join(','),
|
||||||
|
});
|
||||||
|
for (const vid of detailsData?.items || []) if (vid?.id) detailsMap.set(vid.id, vid);
|
||||||
|
} catch (e) {
|
||||||
|
console.warn('[YouTube] details fetch failed, continuing without durations:', e?.message || e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return (lastItems || []).map((item) => {
|
||||||
|
const videoId = item?.id?.videoId;
|
||||||
|
const snippet = item?.snippet || {};
|
||||||
|
const thumb = snippet.thumbnails?.high?.url || snippet.thumbnails?.medium?.url || snippet.thumbnails?.default?.url || undefined;
|
||||||
|
const details = videoId ? detailsMap.get(videoId) : null;
|
||||||
|
const duration = parseISODurationToSeconds(details?.contentDetails?.duration || '');
|
||||||
|
const views = details?.statistics?.viewCount != null ? Number(details.statistics.viewCount) : undefined;
|
||||||
|
const channelId = snippet.channelId || undefined;
|
||||||
|
const embeddable = details?.status ? details.status.embeddable !== false : undefined;
|
||||||
|
return {
|
||||||
|
title: snippet.title || '', id: videoId,
|
||||||
|
url: videoId ? `https://www.youtube.com/watch?v=${videoId}` : undefined,
|
||||||
|
thumbnail: thumb, uploaderName: snippet.channelTitle || undefined, type: 'video',
|
||||||
|
duration: duration > 0 ? duration : undefined, views,
|
||||||
|
publishedAt: snippet.publishedAt || undefined, channelId,
|
||||||
|
channelHandle: snippet.channelTitle || undefined, channelExternalId: channelId,
|
||||||
|
channelUrl: channelId ? `https://www.youtube.com/channel/${channelId}` : undefined, embeddable,
|
||||||
|
};
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
/** @type {{ id: 'yt', label: string, search: (q: string, opts: { limit: number, page?: number, sort?: 'relevance'|'date'|'views' }) => Promise<Suggestion[]> }} */
|
/** @type {{ id: 'yt', label: string, search: (q: string, opts: { limit: number, page?: number, sort?: 'relevance'|'date'|'views' }) => Promise<Suggestion[]> }} */
|
||||||
const handler = {
|
const handler = {
|
||||||
id: 'yt',
|
id: 'yt',
|
||||||
label: 'YouTube',
|
label: 'YouTube',
|
||||||
async search(q, opts) {
|
async search(q, opts) {
|
||||||
const { limit = 10, page = 1, sort = 'relevance' } = opts || {};
|
const { limit = 10, page = 1, sort = 'relevance' } = opts || {};
|
||||||
|
const mode = getSearchMode();
|
||||||
|
const ttl = getScrapeTtlMs();
|
||||||
|
const perPage = Math.min(Math.max(1, Number(limit || 10)), 50);
|
||||||
|
const key = `yt|${hashSearchKey(`${String(q).toLowerCase().trim()}|${perPage}|${page}|${sort}|${mode}`)}`;
|
||||||
|
// 1) mémoire
|
||||||
|
const memHit = memGet(key, ttl);
|
||||||
|
if (memHit) { ytMetrics.scrapeHits++; return memHit; }
|
||||||
|
// 2) SQLite
|
||||||
try {
|
try {
|
||||||
const keys = getYouTubeKeys();
|
const { getCachedYoutubeSearch } = await import('../db.mjs');
|
||||||
if (!keys.length) {
|
const cached = getCachedYoutubeSearch(key);
|
||||||
throw Object.assign(new Error('YOUTUBE_API_KEY not configured'), { ytStatus: 503, code: 'youtube_api_key_unavailable' });
|
if (cached?.items) {
|
||||||
|
ytMetrics.scrapeHits++;
|
||||||
|
memSet(key, cached.items);
|
||||||
|
return cached.items;
|
||||||
}
|
}
|
||||||
|
} catch {}
|
||||||
let order = 'relevance';
|
const persist = async (items, source) => {
|
||||||
if (sort === 'date') order = 'date';
|
// Ne jamais persister un résultat vide : une page vide transitoire
|
||||||
else if (sort === 'views') order = 'viewCount';
|
// (continuation expirée, raté réseau partiel) ne doit pas empoisonner
|
||||||
|
// le cache et bloquer les requêtes suivantes.
|
||||||
// Iterate nextPageToken to reach the requested page (1-based)
|
if (!Array.isArray(items) || items.length === 0) return items;
|
||||||
const perPage = Math.min(Math.max(1, Number(limit || 10)), 50);
|
memSet(key, items);
|
||||||
const targetPage = Math.max(1, Number(page || 1));
|
try {
|
||||||
let pageToken = '';
|
const { setCachedYoutubeSearch, incYoutubeMetrics } = await import('../db.mjs');
|
||||||
let currentPage = 1;
|
setCachedYoutubeSearch(key, q, items, source, ttl);
|
||||||
let lastItems = [];
|
if (source === 'scrape') incYoutubeMetrics({ scrapeCalls: 1 });
|
||||||
|
} catch {}
|
||||||
while (currentPage <= targetPage) {
|
return items;
|
||||||
const params = {
|
};
|
||||||
part: 'snippet',
|
const t0 = Date.now();
|
||||||
q: q,
|
const tryScrape = async () => {
|
||||||
type: 'video',
|
ytMetrics.scrapeCalls++;
|
||||||
maxResults: String(perPage),
|
try { const { incYoutubeMetrics } = await import('../db.mjs'); incYoutubeMetrics({ scrapeCalls: 1 }); } catch {}
|
||||||
order,
|
return searchViaScrape(q, { limit: perPage, page, sort });
|
||||||
// Ne retourner que des vidéos lisibles en embed (évite l'erreur 153 côté player).
|
};
|
||||||
// Les vidéos non-embeddables sont filtrées via le champ status ci-dessous pour /videos.
|
const tryApi = () => searchViaApi(q, { limit: perPage, page, sort });
|
||||||
videoEmbeddable: 'true',
|
const tryInnerTube = () => searchViaInnerTube(q, { limit: perPage, page, sort });
|
||||||
safeSearch: 'moderate',
|
const log = (source, extra = '') => console.log(`[YT search] source=${source} mode=${mode} latency=${Date.now() - t0}ms results=${extra}`);
|
||||||
};
|
try {
|
||||||
if (pageToken) params.pageToken = pageToken;
|
if (mode === 'api-only') {
|
||||||
|
const items = await tryApi();
|
||||||
const data = await ytFetchJson('https://www.googleapis.com/youtube/v3/search', params);
|
log('api', items.length);
|
||||||
|
return persist(items, 'api');
|
||||||
if (currentPage === targetPage) {
|
|
||||||
lastItems = Array.isArray(data.items) ? data.items : [];
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Prepare for next iteration
|
|
||||||
const next = data.nextPageToken;
|
|
||||||
if (!next) {
|
|
||||||
// No more pages; requested page beyond available results
|
|
||||||
lastItems = [];
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
pageToken = String(next);
|
|
||||||
currentPage++;
|
|
||||||
}
|
}
|
||||||
|
if (mode === 'scrape-only') {
|
||||||
const videoIds = (lastItems || [])
|
const items = await tryScrape();
|
||||||
.map(item => item?.id?.videoId)
|
log('scrape', items.length);
|
||||||
.filter(Boolean);
|
return persist(items, 'scrape');
|
||||||
|
}
|
||||||
const detailsMap = new Map();
|
if (mode === 'innertube-only') {
|
||||||
if (videoIds.length > 0) {
|
const items = await tryInnerTube();
|
||||||
|
log('innertube', items.length);
|
||||||
|
return persist(items, 'innertube');
|
||||||
|
}
|
||||||
|
if (mode === 'api-first') {
|
||||||
try {
|
try {
|
||||||
const detailsData = await ytFetchJson('https://www.googleapis.com/youtube/v3/videos', {
|
const items = await tryApi();
|
||||||
part: 'contentDetails,statistics,status',
|
log('api', items.length);
|
||||||
id: videoIds.join(','),
|
return persist(items, 'api');
|
||||||
});
|
} catch (apiErr) {
|
||||||
for (const vid of detailsData?.items || []) {
|
ytMetrics.fallbacks++;
|
||||||
if (vid?.id) detailsMap.set(vid.id, vid);
|
console.warn(`[YT search] api-first fallback to scrape: ${apiErr?.message || apiErr}`);
|
||||||
}
|
const items = await tryScrape();
|
||||||
} catch (e) {
|
log('scrape(fallback)', items.length);
|
||||||
console.warn('[YouTube] details fetch failed, continuing without durations:', e?.message || e);
|
return persist(items, 'scrape');
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
if (mode === 'scrape-first') {
|
||||||
return (lastItems || []).map(item => {
|
// scrape -> API (comportement historique, sans InnerTube)
|
||||||
const videoId = item?.id?.videoId;
|
try {
|
||||||
const snippet = item?.snippet || {};
|
const items = await tryScrape();
|
||||||
const thumb = snippet.thumbnails?.high?.url
|
log('scrape', items.length);
|
||||||
|| snippet.thumbnails?.medium?.url
|
return persist(items, 'scrape');
|
||||||
|| snippet.thumbnails?.default?.url
|
} catch (scrapeErr) {
|
||||||
|| undefined;
|
ytMetrics.scrapeErrors++;
|
||||||
|
ytMetrics.fallbacks++;
|
||||||
const details = videoId ? detailsMap.get(videoId) : null;
|
console.warn(`[YT search] scrape failed (${scrapeErr?.code || 'unknown'}), fallback to api`);
|
||||||
const isoDuration = details?.contentDetails?.duration || '';
|
try {
|
||||||
const duration = parseISODurationToSeconds(isoDuration);
|
const items = await tryApi();
|
||||||
const views = details?.statistics?.viewCount != null ? Number(details.statistics.viewCount) : undefined;
|
log('api(fallback)', items.length);
|
||||||
const channelId = snippet.channelId || undefined;
|
return persist(items, 'api');
|
||||||
const channelHandle = snippet.channelTitle || undefined;
|
} catch (apiErr) {
|
||||||
// status.embeddable === false -> le player renvoie l'erreur 153 ("Video configuration error").
|
const noBinary = scrapeErr?.code === 'yt_scrape_no_binary';
|
||||||
const embeddable = details?.status ? details.status.embeddable !== false : undefined;
|
const noKey = apiErr?.code === 'youtube_api_key_unavailable';
|
||||||
|
if (noBinary && noKey) {
|
||||||
return {
|
const err = new Error('YouTube indisponible : yt-dlp introuvable (définissez YT_DLP_PATH) et aucune clé YOUTUBE_API_KEY configurée.');
|
||||||
title: snippet.title || '',
|
err.ytStatus = 503; err.code = 'youtube_no_source';
|
||||||
id: videoId,
|
throw err;
|
||||||
url: videoId ? `https://www.youtube.com/watch?v=${videoId}` : undefined,
|
}
|
||||||
thumbnail: thumb,
|
console.error(`[YT search] both sources failed. scrape=${scrapeErr?.message} api=${apiErr?.message}`);
|
||||||
uploaderName: snippet.channelTitle || undefined,
|
throw scrapeErr;
|
||||||
type: 'video',
|
}
|
||||||
duration: duration > 0 ? duration : undefined,
|
}
|
||||||
views,
|
}
|
||||||
publishedAt: snippet.publishedAt || undefined,
|
// innertube-first (défaut) : InnerTube -> scrape yt-dlp -> API officielle.
|
||||||
channelId,
|
// Chaque couche ne fait jamais échouer la recherche à elle seule.
|
||||||
channelHandle,
|
const errors = [];
|
||||||
channelExternalId: channelId,
|
try {
|
||||||
channelUrl: channelId ? `https://www.youtube.com/channel/${channelId}` : undefined,
|
const items = await tryInnerTube();
|
||||||
embeddable,
|
log('innertube', items.length);
|
||||||
};
|
return persist(items, 'innertube');
|
||||||
});
|
} catch (itErr) {
|
||||||
|
ytMetrics.innertubeErrors = (ytMetrics.innertubeErrors || 0) + 1;
|
||||||
|
ytMetrics.fallbacks++;
|
||||||
|
errors.push(`innertube=${itErr?.message}`);
|
||||||
|
console.warn(`[YT search] innertube failed (${itErr?.code || 'unknown'}), fallback to scrape`);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
const items = await tryScrape();
|
||||||
|
log('scrape(fallback)', items.length);
|
||||||
|
return persist(items, 'scrape');
|
||||||
|
} catch (scrapeErr) {
|
||||||
|
ytMetrics.scrapeErrors++;
|
||||||
|
ytMetrics.fallbacks++;
|
||||||
|
errors.push(`scrape=${scrapeErr?.message}`);
|
||||||
|
console.warn(`[YT search] scrape failed (${scrapeErr?.code || 'unknown'}), fallback to api`);
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
const items = await tryApi();
|
||||||
|
log('api(fallback)', items.length);
|
||||||
|
return persist(items, 'api');
|
||||||
|
} catch (apiErr) {
|
||||||
|
errors.push(`api=${apiErr?.message}`);
|
||||||
|
const noSource = !getYouTubeKeys().length;
|
||||||
|
if (noSource) {
|
||||||
|
const err = new Error(`YouTube indisponible (${errors.join(' | ').slice(0, 220)}). Vérifiez le réseau/YT_DLP_PATH ou ajoutez YOUTUBE_API_KEYS.`);
|
||||||
|
err.ytStatus = 503; err.code = 'youtube_no_source';
|
||||||
|
throw err;
|
||||||
|
}
|
||||||
|
console.error(`[YT search] all sources failed. ${errors.join(' | ')}`);
|
||||||
|
const err = new Error(errors[0] || 'search_failed');
|
||||||
|
err.ytStatus = 502; err.code = 'youtube_all_sources_failed';
|
||||||
|
throw err;
|
||||||
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
// Log explicite (clé expirée / quota) puis remonte l'erreur pour que
|
|
||||||
// /api/search la reporte dans `errors.yt` au lieu d'un groupe vide silencieux.
|
|
||||||
const status = error?.ytStatus || 500;
|
const status = error?.ytStatus || 500;
|
||||||
console.error(`YouTube search error (status ${status}):`, error?.message || error);
|
console.error(`YouTube search error (status ${status}):`, error?.message || error);
|
||||||
throw error;
|
throw error;
|
||||||
}
|
}
|
||||||
|
},
|
||||||
|
async suggest(q, opts) {
|
||||||
|
const limit = Math.min(20, Math.max(1, Number(opts?.limit || 10)));
|
||||||
|
const query = String(q || '').trim();
|
||||||
|
if (query.length < 2) return [];
|
||||||
|
try {
|
||||||
|
const params = new URLSearchParams({ client: 'youtube', ds: 'yt', hl: 'fr', q: query });
|
||||||
|
const headers = { 'User-Agent': 'Mozilla/5.0' };
|
||||||
|
const proxy = String(process.env.YT_EGRESS_PROXY || '').trim();
|
||||||
|
const fetchOpts = { headers, signal: AbortSignal.timeout(6000) };
|
||||||
|
// Note : suggest reste sans proxy (endpoint Google public léger) sauf si egress configuré côté infra.
|
||||||
|
if (proxy) console.debug?.('[YT suggest] egress proxy configured, direct fetch kept for suggest');
|
||||||
|
const resp = await fetch(`https://suggestqueries.google.com/complete/search?${params.toString()}`, fetchOpts);
|
||||||
|
if (!resp.ok) return [];
|
||||||
|
const text = await resp.text();
|
||||||
|
return parseYoutubeSuggestResponse(text).slice(0, limit);
|
||||||
|
} catch {
|
||||||
|
return [];
|
||||||
|
}
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,130 @@
|
|||||||
|
// Step 15 — Suggest tests: unit parsing/dedup (offline) + API contract (isolated server).
|
||||||
|
// Run with: npm run test:suggest
|
||||||
|
// Scenario coverage (todo Step 15):
|
||||||
|
// - providers=yt returns only yt suggestions
|
||||||
|
// - unknown provider falls back to the full registry
|
||||||
|
// Assertions are on fan-out shape, never on external provider content.
|
||||||
|
|
||||||
|
import fs from 'node:fs';
|
||||||
|
import path from 'node:path';
|
||||||
|
import os from 'node:os';
|
||||||
|
import { spawn } from 'node:child_process';
|
||||||
|
import net from 'node:net';
|
||||||
|
|
||||||
|
function expect(cond, msg) {
|
||||||
|
if (!cond) throw new Error(`Assertion failed: ${msg}`);
|
||||||
|
}
|
||||||
|
function logOk(msg) { console.log(`✓ ${msg}`); }
|
||||||
|
|
||||||
|
// ---------- 1) Offline unit tests: parsers + graceful degradation ----------
|
||||||
|
const yt = (await import('../providers/youtube.mjs')).default;
|
||||||
|
const { parseYoutubeSuggestResponse } = await import('../providers/youtube.mjs');
|
||||||
|
const dm = (await import('../providers/dailymotion.mjs')).default;
|
||||||
|
|
||||||
|
// Bare JSON payload
|
||||||
|
expect(
|
||||||
|
JSON.stringify(parseYoutubeSuggestResponse('["tut",["tutoriel angular","tutoriel android"],[],{}]')) ===
|
||||||
|
JSON.stringify(['tutoriel angular', 'tutoriel android']),
|
||||||
|
'bare JSON suggest payload parses',
|
||||||
|
);
|
||||||
|
logOk('parseYoutubeSuggestResponse bare JSON');
|
||||||
|
|
||||||
|
// window.google.ac.h(...) wrapped payload
|
||||||
|
expect(
|
||||||
|
JSON.stringify(
|
||||||
|
parseYoutubeSuggestResponse('window.google.ac.h(["tut",["tutoriel angular"],[],{}])'),
|
||||||
|
) === JSON.stringify(['tutoriel angular']),
|
||||||
|
'wrapped JSONP suggest payload parses',
|
||||||
|
);
|
||||||
|
logOk('parseYoutubeSuggestResponse wrapped payload');
|
||||||
|
|
||||||
|
// Array-wrapped suggestions (newer format: [text, type, ...])
|
||||||
|
expect(
|
||||||
|
JSON.stringify(parseYoutubeSuggestResponse('["tut",[["tutoriel angular",0],["tutoriel android",0]]]')) ===
|
||||||
|
JSON.stringify(['tutoriel angular', 'tutoriel android']),
|
||||||
|
'array-wrapped suggestions keep text only',
|
||||||
|
);
|
||||||
|
logOk('parseYoutubeSuggestResponse array-wrapped entries');
|
||||||
|
|
||||||
|
// Garbage / empty payloads never throw
|
||||||
|
expect(parseYoutubeSuggestResponse('not json at all').length === 0, 'garbage payload yields []');
|
||||||
|
expect(parseYoutubeSuggestResponse('').length === 0, 'empty payload yields []');
|
||||||
|
logOk('parseYoutubeSuggestResponse graceful degradation');
|
||||||
|
|
||||||
|
// min-length guard: no network below 2 chars (would throw offline otherwise)
|
||||||
|
expect(JSON.stringify(await yt.suggest('a')) === '[]', 'yt.suggest short query returns [] without network');
|
||||||
|
expect(JSON.stringify(await dm.suggest('x')) === '[]', 'dm.suggest short query returns [] without network');
|
||||||
|
logOk('suggest min-length guard (no network)');
|
||||||
|
|
||||||
|
// ---------- 2) API contract against a real isolated server ----------
|
||||||
|
const PORT = await new Promise((resolve) => {
|
||||||
|
const srv = net.createServer();
|
||||||
|
srv.listen(0, '127.0.0.1', () => {
|
||||||
|
const p = srv.address().port;
|
||||||
|
srv.close(() => resolve(p));
|
||||||
|
});
|
||||||
|
});
|
||||||
|
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'newtube-suggest-test-'));
|
||||||
|
const server = spawn(process.execPath, ['./server/index.mjs'], {
|
||||||
|
env: { ...process.env, PORT: String(PORT), NEWTUBE_DB_FILE: path.join(tmpDir, 'suggest.db'), JWT_SECRET: 'suggest-test-secret', NODE_ENV: 'test' },
|
||||||
|
stdio: ['ignore', 'pipe', 'pipe'],
|
||||||
|
cwd: path.resolve(import.meta.dirname, '..', '..'),
|
||||||
|
});
|
||||||
|
let serverLogs = '';
|
||||||
|
server.stdout.on('data', (d) => { serverLogs += d.toString(); });
|
||||||
|
server.stderr.on('data', (d) => { serverLogs += d.toString(); });
|
||||||
|
const baseUrl = `http://127.0.0.1:${PORT}`;
|
||||||
|
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
||||||
|
async function waitForServer(timeoutMs = 20000) {
|
||||||
|
const start = Date.now();
|
||||||
|
while (Date.now() - start < timeoutMs) {
|
||||||
|
try {
|
||||||
|
const res = await fetch(`${baseUrl}/`);
|
||||||
|
if (res.status < 500) return true;
|
||||||
|
} catch {}
|
||||||
|
await sleep(300);
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
let failures = 0;
|
||||||
|
async function scenario(name, fn) {
|
||||||
|
try { await fn(); logOk(name); }
|
||||||
|
catch (e) { failures += 1; console.error(`✗ ${name}`); console.error(e?.stack || String(e)); }
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
const up = await waitForServer();
|
||||||
|
expect(up, 'isolated server starts');
|
||||||
|
|
||||||
|
await scenario('q too short -> 400', async () => {
|
||||||
|
const res = await fetch(`${baseUrl}/api/search/suggest?q=a`);
|
||||||
|
expect(res.status === 400, `expected 400, got ${res.status}`);
|
||||||
|
});
|
||||||
|
|
||||||
|
await scenario('providers=yt returns only yt suggestions', async () => {
|
||||||
|
const res = await fetch(`${baseUrl}/api/search/suggest?q=tutoriel&providers=yt&limit=5`);
|
||||||
|
expect(res.status === 200, `expected 200, got ${res.status}`);
|
||||||
|
const body = await res.json();
|
||||||
|
expect(body.q === 'tutoriel', 'q echoed');
|
||||||
|
expect(JSON.stringify(Object.keys(body.groups)) === JSON.stringify(['yt']), `only yt group (got ${Object.keys(body.groups)})`);
|
||||||
|
expect(Array.isArray(body.groups.yt), 'yt group is an array');
|
||||||
|
expect(body.groups.yt.length <= 5, 'limit respected');
|
||||||
|
});
|
||||||
|
|
||||||
|
await scenario('unknown provider falls back to the full registry', async () => {
|
||||||
|
const res = await fetch(`${baseUrl}/api/search/suggest?q=tutoriel&providers=xx,yy&limit=3`);
|
||||||
|
expect(res.status === 200, `expected 200, got ${res.status}`);
|
||||||
|
const body = await res.json();
|
||||||
|
const keys = Object.keys(body.groups).sort();
|
||||||
|
expect(JSON.stringify(keys) === JSON.stringify(['dm', 'od', 'pt', 'ru', 'tw', 'yt']), `full registry fallback (got ${keys})`);
|
||||||
|
});
|
||||||
|
} finally {
|
||||||
|
server.kill();
|
||||||
|
}
|
||||||
|
|
||||||
|
if (failures > 0) {
|
||||||
|
console.error(`\n${failures} suggest scenario(s) failed. Server logs:\n${serverLogs}`);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
console.log('\nAll suggest tests passed.');
|
||||||
@@ -0,0 +1,212 @@
|
|||||||
|
// Step 16 — Transcript tests (offline fixtures for parsers + API contract).
|
||||||
|
// Run with: npm run test:transcript
|
||||||
|
import { describe, it } from 'node:test';
|
||||||
|
import assert from 'node:assert/strict';
|
||||||
|
import fs from 'node:fs';
|
||||||
|
import path from 'node:path';
|
||||||
|
import os from 'node:os';
|
||||||
|
import { spawn } from 'node:child_process';
|
||||||
|
import net from 'node:net';
|
||||||
|
import {
|
||||||
|
pickTrack,
|
||||||
|
parseJson3,
|
||||||
|
parseVtt,
|
||||||
|
parseTrackText,
|
||||||
|
parseXmlCaptions,
|
||||||
|
orderedTracks,
|
||||||
|
looksLikeHtmlError,
|
||||||
|
capLines,
|
||||||
|
normalizeTranscriptProvider,
|
||||||
|
} from '../transcript.mjs';
|
||||||
|
|
||||||
|
const track = (url, ext = 'vtt') => ({ url, ext });
|
||||||
|
|
||||||
|
describe('pickTrack', () => {
|
||||||
|
it('prefers manual subtitles over automatic captions', () => {
|
||||||
|
const json = {
|
||||||
|
subtitles: { en: [track('https://x/en.vtt')] },
|
||||||
|
automatic_captions: { en: [track('https://x/en.auto.vtt')] },
|
||||||
|
};
|
||||||
|
const { track: t, lang } = pickTrack(json, 'en');
|
||||||
|
assert.equal(t.url, 'https://x/en.vtt');
|
||||||
|
assert.equal(lang, 'en');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('falls back fr -> fr-* -> auto -> first available', () => {
|
||||||
|
const json = {
|
||||||
|
subtitles: { 'fr-CA': [track('https://x/frca.vtt')] },
|
||||||
|
automatic_captions: { en: [track('https://x/en.auto.vtt')] },
|
||||||
|
};
|
||||||
|
assert.equal(pickTrack(json, 'fr').lang, 'fr-CA');
|
||||||
|
// 'de' matches nothing: first available (manual preferred)
|
||||||
|
const fb = pickTrack(json, 'de');
|
||||||
|
assert.equal(fb.lang, 'fr-CA');
|
||||||
|
// auto-only dict
|
||||||
|
const autoOnly = pickTrack({ automatic_captions: { en: [track('https://x/e.vtt')] } }, 'en');
|
||||||
|
assert.equal(autoOnly.track.url, 'https://x/e.vtt');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('returns null track when no subtitles exist', () => {
|
||||||
|
assert.deepEqual(pickTrack({}, 'fr'), { track: null, languages: [], lang: null });
|
||||||
|
assert.deepEqual(pickTrack({ subtitles: {}, automatic_captions: {} }, 'en'), {
|
||||||
|
track: null,
|
||||||
|
languages: [],
|
||||||
|
lang: null,
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it('exposes the language list for the UI selector', () => {
|
||||||
|
const { languages } = pickTrack(
|
||||||
|
{ subtitles: { fr: [track('a')], en: [track('b')] }, automatic_captions: { es: [track('c')] } },
|
||||||
|
'fr',
|
||||||
|
);
|
||||||
|
assert.deepEqual([...languages].sort(), ['en', 'es', 'fr']);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe('parseJson3', () => {
|
||||||
|
it('converts events with segs to normalized lines', () => {
|
||||||
|
const lines = parseJson3({
|
||||||
|
events: [
|
||||||
|
{ tStartMs: 0, dDurationMs: 2500, segs: [{ utf8: 'Bonjour' }] },
|
||||||
|
{ tStartMs: 2500, dDurationMs: 3100, segs: [{ utf8: 'Bien' }, { utf8: 'venue' }] },
|
||||||
|
],
|
||||||
|
});
|
||||||
|
assert.deepEqual(lines, [
|
||||||
|
{ t: 0, dur: 2.5, text: 'Bonjour' },
|
||||||
|
{ t: 2.5, dur: 3.1, text: 'Bienvenue' },
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('drops events without segs or with empty text', () => {
|
||||||
|
const lines = parseJson3({ events: [{ tStartMs: 0 }, { tStartMs: 1, dDurationMs: 1, segs: [{ utf8: ' ' }] }] });
|
||||||
|
assert.deepEqual(lines, []);
|
||||||
|
assert.deepEqual(parseJson3({}), []);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe('parseVtt', () => {
|
||||||
|
const vtt = `WEBVTT
|
||||||
|
|
||||||
|
00:00:00.000 --> 00:00:02.500
|
||||||
|
Bonjour <c.colorE5E5E5>à tous</c>
|
||||||
|
|
||||||
|
12
|
||||||
|
00:00:02.500 --> 00:00:05.600
|
||||||
|
Bienvenue
|
||||||
|
dans cette vidéo
|
||||||
|
`;
|
||||||
|
it('parses cues with identifiers, tags and multi-line content', () => {
|
||||||
|
const lines = parseVtt(vtt);
|
||||||
|
assert.equal(lines.length, 2);
|
||||||
|
assert.equal(lines[0].t, 0);
|
||||||
|
assert.equal(lines[0].dur, 2.5);
|
||||||
|
assert.equal(lines[0].text, 'Bonjour à tous');
|
||||||
|
assert.equal(lines[1].text, 'Bienvenue dans cette vidéo');
|
||||||
|
assert.ok(Math.abs(lines[1].t - 2.5) < 1e-9);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('supports MM:SS.mmm timestamps and decodes entities', () => {
|
||||||
|
const lines = parseVtt('WEBVTT\n\n01:02.000 --> 01:04.500\nFish & chips\n');
|
||||||
|
assert.equal(lines.length, 1);
|
||||||
|
assert.equal(lines[0].t, 62);
|
||||||
|
assert.equal(lines[0].text, 'Fish & chips');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('ignores garbage blocks', () => {
|
||||||
|
assert.deepEqual(parseVtt('WEBVTT\n\nNOTE nothing here\n'), []);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe('orderedTracks + parseTrackText + XML captions', () => {
|
||||||
|
const auto = {
|
||||||
|
fr: [{ url: 'https://x/fr.json3', ext: 'json3' }],
|
||||||
|
en: [{ url: 'https://x/en.json3', ext: 'json3' }],
|
||||||
|
es: [{ url: 'https://x/es.vtt', ext: 'vtt' }],
|
||||||
|
};
|
||||||
|
it('tries requested lang first, then original en', () => {
|
||||||
|
const langs = orderedTracks({ subtitles: {}, automatic_captions: auto }, 'fr').map((o) => o.lang);
|
||||||
|
assert.deepEqual(langs, ['fr', 'en', 'es']);
|
||||||
|
});
|
||||||
|
it('dedupes identical track URLs', () => {
|
||||||
|
const dup = { subtitles: {}, automatic_captions: { en: [{ url: 'https://x/same', ext: 'vtt' }], fr: [{ url: 'https://x/same', ext: 'vtt' }] } };
|
||||||
|
assert.equal(orderedTracks(dup, 'fr').length, 1);
|
||||||
|
});
|
||||||
|
it('detects HTML error pages instead of JSON-parsing them', () => {
|
||||||
|
assert.equal(looksLikeHtmlError('<!DOCTYPE html><html>Sorry</html>'), true);
|
||||||
|
assert.equal(looksLikeHtmlError('WEBVTT\n\n00:00:00.000 --> 1'), false);
|
||||||
|
assert.deepEqual(parseTrackText('<!DOCTYPE html> nope', 'json3'), []);
|
||||||
|
});
|
||||||
|
it('parses srv/ttml XML captions', () => {
|
||||||
|
const srv = parseXmlCaptions('<transcript><p t="0" d="2500">Hello</p></transcript>');
|
||||||
|
assert.equal(srv.length, 1);
|
||||||
|
assert.equal(srv[0].text, 'Hello');
|
||||||
|
const ttml = parseXmlCaptions('<tt><body><div><p begin="00:00:01.000" end="00:00:03.500">Bonjour</p></div></body></tt>');
|
||||||
|
assert.equal(ttml.length, 1);
|
||||||
|
assert.equal(ttml[0].t, 1);
|
||||||
|
});
|
||||||
|
it('parseTrackText handles json3 and vtt payloads', () => {
|
||||||
|
const j = parseTrackText(JSON.stringify({ events: [{ tStartMs: 0, dDurationMs: 1000, segs: [{ utf8: 'Hi' }] }] }), 'json3');
|
||||||
|
assert.equal(j.length, 1);
|
||||||
|
const v = parseTrackText('WEBVTT\n\n00:00:00.000 --> 00:00:01.000\nHello\n', 'vtt');
|
||||||
|
assert.equal(v.length, 1);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe('capLines + normalizeTranscriptProvider', () => { it('truncates very long transcripts', () => {
|
||||||
|
const lines = Array.from({ length: 10 }, (_, i) => ({ t: i, dur: 1, text: `l${i}` }));
|
||||||
|
assert.equal(capLines(lines, 3).length, 3);
|
||||||
|
});
|
||||||
|
it('maps short and long provider ids', () => {
|
||||||
|
assert.equal(normalizeTranscriptProvider('yt'), 'youtube');
|
||||||
|
assert.equal(normalizeTranscriptProvider('youtube'), 'youtube');
|
||||||
|
assert.equal(normalizeTranscriptProvider('dm'), 'dailymotion');
|
||||||
|
assert.equal(normalizeTranscriptProvider('pt'), 'peertube');
|
||||||
|
assert.equal(normalizeTranscriptProvider('xx'), null);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe('API contract', () => {
|
||||||
|
it('rejects unknown providers with 400 { available: false }', async () => {
|
||||||
|
const port = await new Promise((resolve) => {
|
||||||
|
const srv = net.createServer();
|
||||||
|
srv.listen(0, '127.0.0.1', () => {
|
||||||
|
const p = srv.address().port;
|
||||||
|
srv.close(() => resolve(p));
|
||||||
|
});
|
||||||
|
});
|
||||||
|
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'newtube-transcript-test-'));
|
||||||
|
const server = spawn(process.execPath, ['./server/index.mjs'], {
|
||||||
|
env: {
|
||||||
|
...process.env,
|
||||||
|
PORT: String(port),
|
||||||
|
NEWTUBE_DB_FILE: path.join(tmpDir, 'transcript.db'),
|
||||||
|
JWT_SECRET: 'transcript-test-secret',
|
||||||
|
NODE_ENV: 'test',
|
||||||
|
},
|
||||||
|
stdio: ['ignore', 'pipe', 'pipe'],
|
||||||
|
cwd: path.resolve(import.meta.dirname, '..', '..'),
|
||||||
|
});
|
||||||
|
const baseUrl = `http://127.0.0.1:${port}`;
|
||||||
|
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
||||||
|
try {
|
||||||
|
let up = false;
|
||||||
|
const start = Date.now();
|
||||||
|
while (Date.now() - start < 20000) {
|
||||||
|
try {
|
||||||
|
const res = await fetch(`${baseUrl}/`);
|
||||||
|
if (res.status < 500) { up = true; break; }
|
||||||
|
} catch {}
|
||||||
|
await sleep(300);
|
||||||
|
}
|
||||||
|
assert.ok(up, 'isolated server starts');
|
||||||
|
const res = await fetch(`${baseUrl}/api/transcript/xx/abc123?lang=fr`);
|
||||||
|
assert.equal(res.status, 400);
|
||||||
|
const body = await res.json();
|
||||||
|
assert.equal(body.available, false);
|
||||||
|
assert.ok(typeof body.error === 'string');
|
||||||
|
} finally {
|
||||||
|
server.kill();
|
||||||
|
}
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -0,0 +1,131 @@
|
|||||||
|
// Step 18 — tests offline InnerTube (AUCUN réseau : mappers purs + modes).
|
||||||
|
import assert from 'node:assert/strict';
|
||||||
|
import { mapVideoNode, mapLockupView, mapNodes, mapCaptionTracks, parseViewsText, parseDurationLabel } from '../providers/youtube-innertube.mjs';
|
||||||
|
import { pickTrack, orderedTracks } from '../transcript.mjs';
|
||||||
|
import { getSearchMode, YT_SEARCH_MODES } from '../providers/youtube-common.mjs';
|
||||||
|
|
||||||
|
function ok(msg) { console.log(`✓ ${msg}`); }
|
||||||
|
|
||||||
|
// 1) parseViewsText FR/EN
|
||||||
|
{
|
||||||
|
assert.equal(parseViewsText('81 973 vues'), 81973);
|
||||||
|
assert.equal(parseViewsText('1,2 M vues'), 1200000);
|
||||||
|
assert.equal(parseViewsText('57K'), 57000);
|
||||||
|
assert.equal(parseViewsText('3 k vues'), 3000);
|
||||||
|
assert.equal(parseViewsText(''), undefined);
|
||||||
|
assert.equal(parseViewsText(null), undefined);
|
||||||
|
ok('parseViewsText FR/EN');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 1b) parseDurationLabel (a11y)
|
||||||
|
{
|
||||||
|
assert.equal(parseDurationLabel("What's new in Angular 18 2 minutes, 27 seconds"), 147);
|
||||||
|
assert.equal(parseDurationLabel('Live 1 heure, 5 minutes'), 3900);
|
||||||
|
assert.equal(parseDurationLabel('nonsense'), undefined);
|
||||||
|
ok('parseDurationLabel');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 2) mapVideoNode type Video (forme réelle youtubei.js v18)
|
||||||
|
{
|
||||||
|
const m = mapVideoNode({
|
||||||
|
type: 'Video', id: 'abc123',
|
||||||
|
title: { text: 'Demo' },
|
||||||
|
duration: { text: '15:22', seconds: 922 },
|
||||||
|
thumbnails: [{ url: 'http://t/low.jpg' }, { url: 'http://t/high.jpg' }],
|
||||||
|
author: { id: 'UC999', name: 'Chaine' },
|
||||||
|
published: { text: 'il y a 2 ans' },
|
||||||
|
view_count: { text: '81 973 vues' },
|
||||||
|
badges: [{ label: 'Sous-titres' }],
|
||||||
|
is_live: false,
|
||||||
|
});
|
||||||
|
assert.equal(m.id, 'abc123');
|
||||||
|
assert.equal(m.url, 'https://www.youtube.com/watch?v=abc123');
|
||||||
|
assert.equal(m.thumbnail, 'http://t/high.jpg');
|
||||||
|
assert.equal(m.duration, 922);
|
||||||
|
assert.equal(m.views, 81973);
|
||||||
|
assert.equal(m.channelId, 'UC999');
|
||||||
|
assert.deepEqual(m.badges, ['Sous-titres']);
|
||||||
|
assert.equal(m.isShort, undefined);
|
||||||
|
ok('mapVideoNode Video');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 3) Shorts/Reel -> isShort ; live -> isLive ; bruit filtré
|
||||||
|
{
|
||||||
|
const s = mapVideoNode({ type: 'ShortsLockupView', id: 'sh1', title: { text: 'Short' }, thumbnails: [{ url: 'u' }] });
|
||||||
|
assert.equal(s.isShort, true);
|
||||||
|
const l = mapVideoNode({ type: 'Video', id: 'lv1', title: { text: 'Live' }, thumbnails: [{ url: 'u' }], is_live: true });
|
||||||
|
assert.equal(l.isLive, true);
|
||||||
|
assert.equal(mapVideoNode({ type: 'Playlist', id: 'PL1', title: { text: 'PL' } }), null);
|
||||||
|
assert.equal(mapVideoNode({ type: 'Channel', id: 'UC1', title: { text: 'C' } }), null);
|
||||||
|
assert.equal(mapVideoNode(null), null);
|
||||||
|
assert.equal(mapVideoNode({ type: 'Video', title: { text: 'no id' } }), null);
|
||||||
|
ok('mapVideoNode shorts/live/filtrage');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 4) mapNodes déduplique
|
||||||
|
{
|
||||||
|
const mk = (id) => ({ type: 'Video', id, title: { text: id }, thumbnails: [{ url: 'u' }] });
|
||||||
|
const out = mapNodes([mk('a'), mk('a'), mk('b'), null]);
|
||||||
|
assert.deepEqual(out.map((x) => x.id), ['a', 'b']);
|
||||||
|
ok('mapNodes dédup');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 4b) mapLockupView (renderer moderne du watch-next façon SmartTube)
|
||||||
|
{
|
||||||
|
const m = mapLockupView({
|
||||||
|
type: 'LockupView', content_id: '0wIL6d5TxGc', content_type: 'VIDEO',
|
||||||
|
content_image: { image: [{ url: 'http://t/small.jpg', width: 168 }, { url: 'http://t/big.jpg', width: 336 }] },
|
||||||
|
metadata: {
|
||||||
|
title: { text: "What's new" },
|
||||||
|
metadata: { metadata_rows: [
|
||||||
|
{ metadata_parts: [{ text: { text: 'Rainer Hahnekamp' } }] },
|
||||||
|
{ metadata_parts: [{ text: { text: '57K' } }] },
|
||||||
|
] },
|
||||||
|
image: { a11y_label: 'Go to channel Rainer Hahnekamp' },
|
||||||
|
},
|
||||||
|
renderer_context: { accessibility_context: { label: "What's new 2 minutes, 27 seconds" } },
|
||||||
|
});
|
||||||
|
assert.equal(m.id, '0wIL6d5TxGc');
|
||||||
|
assert.equal(m.url, 'https://www.youtube.com/watch?v=0wIL6d5TxGc');
|
||||||
|
assert.equal(m.thumbnail, 'http://t/big.jpg');
|
||||||
|
assert.equal(m.uploaderName, 'Rainer Hahnekamp');
|
||||||
|
assert.equal(m.views, 57000);
|
||||||
|
assert.equal(m.duration, 147);
|
||||||
|
assert.equal(mapVideoNode({ type: 'LockupView', content_id: 'x', content_type: 'VIDEO', metadata: { title: { text: 'T' } } }).id, 'x');
|
||||||
|
assert.equal(mapLockupView({ type: 'LockupView', content_type: 'PLAYLIST', content_id: 'PL1', metadata: { title: { text: 'PL' } } }), null);
|
||||||
|
ok('mapLockupView watch-next');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 5) modes innertube connus
|
||||||
|
{
|
||||||
|
assert.ok(YT_SEARCH_MODES.includes('innertube-first'));
|
||||||
|
assert.ok(YT_SEARCH_MODES.includes('innertube-only'));
|
||||||
|
const prev = process.env.YT_SEARCH_MODE;
|
||||||
|
process.env.YT_SEARCH_MODE = 'innertube-only';
|
||||||
|
assert.equal(getSearchMode(), 'innertube-only');
|
||||||
|
if (prev === undefined) delete process.env.YT_SEARCH_MODE; else process.env.YT_SEARCH_MODE = prev;
|
||||||
|
ok('modes innertube-first/only');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 6) mapCaptionTracks : forme InnerTube -> pseudo dump yt-dlp, interop pickTrack
|
||||||
|
{
|
||||||
|
const cap = mapCaptionTracks([
|
||||||
|
{ base_url: 'https://x/t?caps=asr&x=1', language_code: 'en', kind: 'asr', name: { text: 'English (auto-generated)' } },
|
||||||
|
{ base_url: 'https://x/t?caps=m&x=2', language_code: 'fr', name: { text: 'Français' } },
|
||||||
|
{ base_url: '', language_code: 'de', name: 'Deutsch' },
|
||||||
|
]);
|
||||||
|
assert.deepEqual(cap.languages.sort(), ['en', 'fr']);
|
||||||
|
assert.equal(cap.trackCount, 2);
|
||||||
|
assert.match(cap.automatic_captions.en[0].url, /fmt=json3/);
|
||||||
|
assert.equal(cap.automatic_captions.en[0].ext, 'json3');
|
||||||
|
assert.equal(cap.subtitles.fr[0].name, 'Français');
|
||||||
|
// Interop : le pseudo dump passe dans pickTrack/orderedTracks sans modification
|
||||||
|
const picked = pickTrack({ subtitles: cap.subtitles, automatic_captions: cap.automatic_captions }, 'fr');
|
||||||
|
assert.equal(picked.lang, 'fr');
|
||||||
|
assert.ok(picked.track.url.includes('fmt=json3'));
|
||||||
|
const ordered = orderedTracks({ subtitles: cap.subtitles, automatic_captions: cap.automatic_captions }, 'de');
|
||||||
|
assert.ok(ordered.length >= 2);
|
||||||
|
ok('mapCaptionTracks + interop pickTrack/orderedTracks');
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log('youtube-innertube tests passed');
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
// Step 17 — tests offline : parsers scrape + dispatcher (AUCUN réseau, AUCUNE clé).
|
||||||
|
import assert from 'node:assert/strict';
|
||||||
|
import { mapFlatEntry, parseFlatPlaylistJson, classifyScrapeError } from '../providers/youtube-scrape.mjs';
|
||||||
|
import { getSearchMode, hashSearchKey, buildYtDlpExtraArgs, resolveYtDlpBin, resetYtDlpBinCache } from '../providers/youtube-common.mjs';
|
||||||
|
|
||||||
|
function ok(msg) { console.log(`✓ ${msg}`); }
|
||||||
|
|
||||||
|
// 1) mapFlatEntry
|
||||||
|
{
|
||||||
|
const s = mapFlatEntry({
|
||||||
|
id: 'dQw4w9WgXcQ', title: 'Demo', channel_id: 'UC123', channel: 'Chaine',
|
||||||
|
duration: 212, view_count: 1000, timestamp: 1700000000,
|
||||||
|
thumbnails: [{ url: 'http://t/low.jpg', height: 90 }, { url: 'http://t/high.jpg', height: 720 }],
|
||||||
|
});
|
||||||
|
assert.equal(s.id, 'dQw4w9WgXcQ');
|
||||||
|
assert.equal(s.url, 'https://www.youtube.com/watch?v=dQw4w9WgXcQ');
|
||||||
|
assert.equal(s.thumbnail, 'http://t/high.jpg');
|
||||||
|
assert.equal(s.duration, 212);
|
||||||
|
ok('mapFlatEntry maps flat entry to Suggestion');
|
||||||
|
}
|
||||||
|
assert.equal(mapFlatEntry(null), null);
|
||||||
|
assert.equal(mapFlatEntry({}), null);
|
||||||
|
ok('mapFlatEntry null-safe');
|
||||||
|
|
||||||
|
// 2) parseFlatPlaylistJson : objet + NDJSON + vide
|
||||||
|
{
|
||||||
|
const raw = JSON.stringify({ entries: [{ id: 'a', title: 'A' }, { id: '', title: 'x' }, null] });
|
||||||
|
assert.equal(parseFlatPlaylistJson(raw).length, 1);
|
||||||
|
ok('parseFlatPlaylistJson objet .entries');
|
||||||
|
}
|
||||||
|
{
|
||||||
|
const raw = `{"id":"a","title":"A"}\n{"id":"b","title":"B"}\nnot-json\n`;
|
||||||
|
assert.equal(parseFlatPlaylistJson(raw).length, 2);
|
||||||
|
ok('parseFlatPlaylistJson NDJSON tolerant');
|
||||||
|
}
|
||||||
|
assert.deepEqual(parseFlatPlaylistJson(''), []);
|
||||||
|
assert.deepEqual(parseFlatPlaylistJson('{{{'), []);
|
||||||
|
ok('parseFlatPlaylistJson vide/invalide -> []');
|
||||||
|
|
||||||
|
// 3) getSearchMode défaut innertube-first + normalisation
|
||||||
|
{
|
||||||
|
const prev = process.env.YT_SEARCH_MODE;
|
||||||
|
delete process.env.YT_SEARCH_MODE;
|
||||||
|
assert.equal(getSearchMode(), 'innertube-first');
|
||||||
|
process.env.YT_SEARCH_MODE = 'API-FIRST';
|
||||||
|
assert.equal(getSearchMode(), 'api-first');
|
||||||
|
process.env.YT_SEARCH_MODE = 'scrape-first';
|
||||||
|
assert.equal(getSearchMode(), 'scrape-first');
|
||||||
|
process.env.YT_SEARCH_MODE = 'nonsense';
|
||||||
|
assert.equal(getSearchMode(), 'innertube-first');
|
||||||
|
if (prev === undefined) delete process.env.YT_SEARCH_MODE; else process.env.YT_SEARCH_MODE = prev;
|
||||||
|
ok('getSearchMode défaut + normalisation');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 4) hash stable + extra args jamais de secret en clair autre que proxy path
|
||||||
|
{
|
||||||
|
assert.equal(hashSearchKey('a'), hashSearchKey('a'));
|
||||||
|
assert.notEqual(hashSearchKey('a'), hashSearchKey('b'));
|
||||||
|
const args = buildYtDlpExtraArgs();
|
||||||
|
assert.ok(Array.isArray(args));
|
||||||
|
ok('hashSearchKey + buildYtDlpExtraArgs');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 5) dispatcher scrape-only sans clé : ne doit PAS exiger YOUTUBE_API_KEY au import,
|
||||||
|
// et searchViaScrape vide sur query courte
|
||||||
|
{
|
||||||
|
const { searchViaScrape } = await import('../providers/youtube-scrape.mjs');
|
||||||
|
assert.deepEqual(await searchViaScrape('a'), []);
|
||||||
|
ok('searchViaScrape query courte -> [] sans réseau');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 6) classifyScrapeError : ENOENT -> yt_scrape_no_binary (message actionnable, jamais de spawn brut en UI)
|
||||||
|
{
|
||||||
|
const e = classifyScrapeError(Object.assign(new Error('spawn yt-dlp ENOENT'), { code: 'ENOENT', errno: 'ENOENT' }));
|
||||||
|
assert.equal(e.code, 'yt_scrape_no_binary');
|
||||||
|
assert.equal(e.ytStatus, 503);
|
||||||
|
assert.match(e.message, /YT_DLP_PATH/);
|
||||||
|
const passthrough = classifyScrapeError(Object.assign(new Error('x'), { code: 'yt_scrape_no_binary' }));
|
||||||
|
assert.equal(passthrough.code, 'yt_scrape_no_binary');
|
||||||
|
ok('classifyScrapeError ENOENT -> yt_scrape_no_binary');
|
||||||
|
}
|
||||||
|
|
||||||
|
// 7) resolveYtDlpBin : YT_DLP_PATH inexistant -> fallback PATH/bundled (jamais de throw si un binaire existe)
|
||||||
|
{
|
||||||
|
resetYtDlpBinCache();
|
||||||
|
const prev = process.env.YT_DLP_PATH;
|
||||||
|
process.env.YT_DLP_PATH = '/nonexistent/yt-dlp';
|
||||||
|
try {
|
||||||
|
const bin = await resolveYtDlpBin();
|
||||||
|
assert.ok(typeof bin === 'string' && bin.length > 0);
|
||||||
|
ok(`resolveYtDlpBin fallback -> ${bin}`);
|
||||||
|
} finally {
|
||||||
|
resetYtDlpBinCache();
|
||||||
|
if (prev === undefined) delete process.env.YT_DLP_PATH; else process.env.YT_DLP_PATH = prev;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log('youtube-scrape tests passed');
|
||||||
@@ -0,0 +1,309 @@
|
|||||||
|
// Step 16 — Pure transcript helpers (no network, no DB).
|
||||||
|
// Data source: `yt-dlp --dump-single-json --skip-download` exposes
|
||||||
|
// `subtitles` (manual) and `automatic_captions` (auto-generated).
|
||||||
|
// Track entries are arrays of { url, ext, name } (or single objects).
|
||||||
|
|
||||||
|
/**
|
||||||
|
* @typedef {{ t: number, dur: number, text: string }} TranscriptLine
|
||||||
|
*/
|
||||||
|
|
||||||
|
const MAX_LINES_DEFAULT = Number(process.env.TRANSCRIPT_MAX_LINES || 5000);
|
||||||
|
|
||||||
|
/** Normalize a language code for comparison (lowercase, `_` -> `-`). */
|
||||||
|
function normLang(code) {
|
||||||
|
return String(code || '').trim().toLowerCase().replace(/_/g, '-');
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Pick the first usable track object from an array-or-single entry. */
|
||||||
|
function firstTrack(entry) {
|
||||||
|
if (!entry) return null;
|
||||||
|
const list = Array.isArray(entry) ? entry : [entry];
|
||||||
|
return list.find((t) => t && typeof t.url === 'string' && t.url) || null;
|
||||||
|
}
|
||||||
|
|
||||||
|
function trackExtOf(track) {
|
||||||
|
const ext = String(track?.ext || '').toLowerCase();
|
||||||
|
if (ext) return ext;
|
||||||
|
try {
|
||||||
|
const u = new URL(String(track?.url || ''));
|
||||||
|
const m = /\.([a-z0-9]+)(?:[?#]|$)/i.exec(u.pathname || '');
|
||||||
|
if (m) return m[1].toLowerCase();
|
||||||
|
} catch {}
|
||||||
|
return '';
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Find a language key in `dict` matching `code` exactly or by prefix (`fr` -> `fr-*`). */
|
||||||
|
function findLangKey(dict, code) {
|
||||||
|
const want = normLang(code);
|
||||||
|
if (!want || !dict) return null;
|
||||||
|
const keys = Object.keys(dict);
|
||||||
|
const exact = keys.find((k) => normLang(k) === want);
|
||||||
|
if (exact) return exact;
|
||||||
|
// `fr-ca` requested, `fr` available (and vice versa): match on primary subtag
|
||||||
|
const primary = want.split('-')[0];
|
||||||
|
return keys.find((k) => normLang(k).split('-')[0] === primary) || null;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Select the best subtitle track.
|
||||||
|
* Priority: manual exact > manual prefix/primary > auto exact > auto prefix/primary > first available.
|
||||||
|
* @param {any} json yt-dlp dump-single-json
|
||||||
|
* @param {string} [lang]
|
||||||
|
* @returns {{ track: any|null, languages: string[], lang: string|null }}
|
||||||
|
*/
|
||||||
|
export function pickTrack(json, lang = 'fr') {
|
||||||
|
const manual = (json && json.subtitles) || {};
|
||||||
|
const auto = (json && json.automatic_captions) || {};
|
||||||
|
const manualKeys = Object.keys(manual);
|
||||||
|
const autoKeys = Object.keys(auto);
|
||||||
|
const languages = Array.from(new Set([...manualKeys, ...autoKeys]));
|
||||||
|
|
||||||
|
if (languages.length === 0) return { track: null, languages: [], lang: null };
|
||||||
|
|
||||||
|
const manualKey = findLangKey(manual, lang);
|
||||||
|
if (manualKey) {
|
||||||
|
const track = firstTrack(manual[manualKey]);
|
||||||
|
if (track) return { track, languages, lang: manualKey };
|
||||||
|
}
|
||||||
|
const autoKey = findLangKey(auto, lang);
|
||||||
|
if (autoKey) {
|
||||||
|
const track = firstTrack(auto[autoKey]);
|
||||||
|
if (track) return { track, languages, lang: autoKey };
|
||||||
|
}
|
||||||
|
// Fallback: first available track (manual preferred)
|
||||||
|
for (const key of manualKeys) {
|
||||||
|
const track = firstTrack(manual[key]);
|
||||||
|
if (track) return { track, languages, lang: key };
|
||||||
|
}
|
||||||
|
for (const key of autoKeys) {
|
||||||
|
const track = firstTrack(auto[key]);
|
||||||
|
if (track) return { track, languages, lang: key };
|
||||||
|
}
|
||||||
|
return { track: null, languages, lang: null };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Order candidate tracks to try when the preferred one cannot be fetched.
|
||||||
|
* YouTube auto-translated tracks (`tlang=xx`) are often rate-limited (HTTP 429
|
||||||
|
* or "Sorry" HTML pages) while the original language still works: always try
|
||||||
|
* the requested language first, then the original (`en`), then everything
|
||||||
|
* else (manual preferred). De-duplicated by URL.
|
||||||
|
* @param {any} json yt-dlp dump-single-json
|
||||||
|
* @param {string} [lang]
|
||||||
|
* @returns {Array<{ track: any, lang: string|null }>}
|
||||||
|
*/
|
||||||
|
export function orderedTracks(json, lang = 'fr') {
|
||||||
|
const manual = (json && json.subtitles) || {};
|
||||||
|
const auto = (json && json.automatic_captions) || {};
|
||||||
|
const seen = new Set();
|
||||||
|
const out = [];
|
||||||
|
const push = (key, dict) => {
|
||||||
|
if (key == null) return;
|
||||||
|
const t = firstTrack(dict[key]);
|
||||||
|
if (!t || !t.url || seen.has(String(t.url))) return;
|
||||||
|
seen.add(String(t.url));
|
||||||
|
out.push({ track: t, lang: key });
|
||||||
|
};
|
||||||
|
const primary = pickTrack(json, lang);
|
||||||
|
if (primary.track) {
|
||||||
|
seen.add(String(primary.track.url));
|
||||||
|
out.push({ track: primary.track, lang: primary.lang });
|
||||||
|
}
|
||||||
|
// Original language first fallback (avoids translated-track rate limits)
|
||||||
|
const reqPrimary = String(lang || '').toLowerCase().split('-')[0];
|
||||||
|
for (const fallbackLang of ['en', 'fr']) {
|
||||||
|
if (fallbackLang === reqPrimary) continue;
|
||||||
|
const k = findLangKey(manual, fallbackLang);
|
||||||
|
if (k) push(k, manual);
|
||||||
|
const ka = findLangKey(auto, fallbackLang);
|
||||||
|
if (ka) push(ka, auto);
|
||||||
|
}
|
||||||
|
for (const key of Object.keys(manual)) push(key, manual);
|
||||||
|
for (const key of Object.keys(auto)) push(key, auto);
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse a YouTube `json3` timedtext payload into normalized lines.
|
||||||
|
* @param {any} data parsed JSON
|
||||||
|
* @returns {TranscriptLine[]}
|
||||||
|
*/
|
||||||
|
export function parseJson3(data) {
|
||||||
|
const events = (data && data.events) || [];
|
||||||
|
if (!Array.isArray(events)) return [];
|
||||||
|
const out = [];
|
||||||
|
for (const e of events) {
|
||||||
|
if (!e || !Array.isArray(e.segs)) continue;
|
||||||
|
const text = e.segs.map((s) => String(s?.utf8 ?? '')).join('').replace(/\n+/g, ' ').trim();
|
||||||
|
if (!text) continue;
|
||||||
|
out.push({
|
||||||
|
t: Number(e.tStartMs || 0) / 1000,
|
||||||
|
dur: Number(e.dDurationMs || 0) / 1000,
|
||||||
|
text,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
return capLines(out);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Decode a handful of HTML entities found in VTT payloads. */
|
||||||
|
function decodeEntities(s) {
|
||||||
|
return String(s || '')
|
||||||
|
.replace(/&/g, '&')
|
||||||
|
.replace(/</g, '<')
|
||||||
|
.replace(/>/g, '>')
|
||||||
|
.replace(/"/g, '"')
|
||||||
|
.replace(/'/g, "'")
|
||||||
|
.replace(/ /g, ' ');
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Strip XML/HTML tags and decode entities to plain cue text. */
|
||||||
|
function xmlTextToPlain(s) {
|
||||||
|
return decodeEntities(String(s || '').replace(/<[^>]*>/g, ' ').replace(/\s+/g, ' ').trim()).trim();
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseTimeAttrToSeconds(v) {
|
||||||
|
if (v == null || v === '') return null;
|
||||||
|
const s = String(v).trim();
|
||||||
|
if (!s) return null;
|
||||||
|
if (/^\d+(\.\d+)?$/.test(s)) return Number(s);
|
||||||
|
const m = /(?:(\d+):)?([0-5]?\d):([0-5]\d)(?:\.(\d{1,3}))?/.exec(s);
|
||||||
|
if (!m) return null;
|
||||||
|
return Number(m[1] || 0) * 3600 + Number(m[2]) * 60 + Number(m[3]) + Number((m[4] || '0').padEnd(3, '0')) / 1000;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse YouTube XML caption formats (`srv1`/`srv2`/`srv3` `<p>` cues,
|
||||||
|
* `ttml` `<p begin= end=>` cues) into normalized lines.
|
||||||
|
* @param {string} text raw XML
|
||||||
|
* @returns {TranscriptLine[]}
|
||||||
|
*/
|
||||||
|
export function parseXmlCaptions(text) {
|
||||||
|
const out = [];
|
||||||
|
const src = String(text || '');
|
||||||
|
if (!src || !/<(p|text|tt|transcript)\b/i.test(src)) return capLines(out);
|
||||||
|
// YouTube srv*: <p t="0" d="2500">Hello <s>world</s></p>
|
||||||
|
const srvRe = /<p\b[^>]*?(?:t|start)="([^"]*)"[^>]*?(?:d|dur)="([^"]*)"[^>]*>([\s\S]*?)<\/p\s*>/gi;
|
||||||
|
// Generic ttml: <p begin="00:00:00.000" end="00:00:02.500">...</p>
|
||||||
|
const ttmlRe = /<p\b[^>]*?begin="([^"]*)"[^>]*?end="([^"]*)"[^>]*>([\s\S]*?)<\/p\s*>/gi;
|
||||||
|
let m;
|
||||||
|
while ((m = srvRe.exec(src)) !== null) {
|
||||||
|
const start = Number(m[1]) / 1000;
|
||||||
|
const dur = Number(m[2]) / 1000;
|
||||||
|
const cleaned = xmlTextToPlain(m[3]);
|
||||||
|
if (!cleaned || !Number.isFinite(start)) continue;
|
||||||
|
out.push({ t: start, dur: Number.isFinite(dur) && dur > 0 ? dur : 0, text: cleaned });
|
||||||
|
}
|
||||||
|
if (out.length === 0) {
|
||||||
|
while ((m = ttmlRe.exec(src)) !== null) {
|
||||||
|
const start = parseTimeAttrToSeconds(m[1]);
|
||||||
|
const end = parseTimeAttrToSeconds(m[2]);
|
||||||
|
const cleaned = xmlTextToPlain(m[3]);
|
||||||
|
if (!cleaned || start == null || end == null) continue;
|
||||||
|
out.push({ t: start, dur: Math.max(0, end - start), text: cleaned });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// Plain <text start="2.5" dur="2.0">Fallback</text> variant
|
||||||
|
if (out.length === 0) {
|
||||||
|
const textRe = /<text\b[^>]*?start="([^"]*)"[^>]*?(?:dur="([^"]*)")?[^>]*>([\s\S]*?)<\/text\s*>/gi;
|
||||||
|
while ((m = textRe.exec(src)) !== null) {
|
||||||
|
const start = Number(m[1]);
|
||||||
|
const dur = Number(m[2] || 0);
|
||||||
|
const cleaned = xmlTextToPlain(m[3]);
|
||||||
|
if (!cleaned || !Number.isFinite(start)) continue;
|
||||||
|
out.push({ t: start, dur: dur > 0 ? dur : 0, text: cleaned });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return capLines(out);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Detect YouTube "Sorry / automated queries" style HTML error pages. */
|
||||||
|
export function looksLikeHtmlError(text) {
|
||||||
|
const s = String(text || '').trimStart().slice(0, 512).toLowerCase();
|
||||||
|
return s.startsWith('<!doctype') || s.startsWith('<html');
|
||||||
|
}
|
||||||
|
|
||||||
|
function vttTimestampToSeconds(h, m, s, ms) {
|
||||||
|
return Number(h) * 3600 + Number(m) * 60 + Number(s) + Number(ms) / 1000;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse a WebVTT payload into normalized lines.
|
||||||
|
* Handles cue identifiers, multi-line cues and inline tags (`<c>`, `<i>`, timestamps).
|
||||||
|
* @param {string} text raw VTT
|
||||||
|
* @returns {TranscriptLine[]}
|
||||||
|
*/
|
||||||
|
export function parseVtt(text) {
|
||||||
|
if (looksLikeHtmlError(text)) return [];
|
||||||
|
const out = [];
|
||||||
|
const blocks = String(text || '').replace(/\r\n/g, '\n').split(/\n{2,}/);
|
||||||
|
const tsRe = /(?:(\d{2,}):)?([0-5]?\d):([0-5]\d)\.(\d{3})\s*-->\s*(?:(\d{2,}):)?([0-5]?\d):([0-5]\d)\.(\d{3})/;
|
||||||
|
for (const block of blocks) {
|
||||||
|
const lines = String(block || '').split('\n');
|
||||||
|
const timingIdx = lines.findIndex((l) => l.includes('-->'));
|
||||||
|
if (timingIdx === -1) continue;
|
||||||
|
const m = tsRe.exec(lines[timingIdx]);
|
||||||
|
if (!m) continue;
|
||||||
|
const start = vttTimestampToSeconds(m[1] || '0', m[2], m[3], m[4]);
|
||||||
|
const end = vttTimestampToSeconds(m[5] || '0', m[6], m[7], m[8]);
|
||||||
|
const content = lines
|
||||||
|
.slice(timingIdx + 1)
|
||||||
|
.join(' ')
|
||||||
|
// strip inline VTT tags (<c.colorE5E5E5>, <i>, <00:00:01.000>, ...)
|
||||||
|
.replace(/<[^>]*>/g, ' ')
|
||||||
|
.replace(/\s+/g, ' ')
|
||||||
|
.trim();
|
||||||
|
const decoded = decodeEntities(content).trim();
|
||||||
|
// Skip WEBVTT header leftovers and empty cues
|
||||||
|
if (!decoded || /^WEBVTT/i.test(decoded)) continue;
|
||||||
|
out.push({ t: start, dur: Math.max(0, end - start), text: decoded });
|
||||||
|
}
|
||||||
|
return capLines(out);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Bound response volume for very long transcripts (Phase 1 decision: truncate). */
|
||||||
|
export function capLines(lines, max = MAX_LINES_DEFAULT) {
|
||||||
|
const list = Array.isArray(lines) ? lines : [];
|
||||||
|
const n = Math.max(1, Number(max || MAX_LINES_DEFAULT));
|
||||||
|
return list.length > n ? list.slice(0, n) : list;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse a downloaded timedtext payload regardless of its format.
|
||||||
|
* Tries `json3`, then `vtt`, then XML captions (`srv*`/`ttml`).
|
||||||
|
* HTML error pages always yield `[]`.
|
||||||
|
*/
|
||||||
|
export function parseTrackText(text, ext = '') {
|
||||||
|
if (looksLikeHtmlError(text)) return [];
|
||||||
|
const e = String(ext || '').toLowerCase();
|
||||||
|
const asJson3 = () => {
|
||||||
|
try { return parseJson3(JSON.parse(String(text))); } catch { return []; }
|
||||||
|
};
|
||||||
|
if (e === 'json3') {
|
||||||
|
const lines = asJson3();
|
||||||
|
if (lines.length) return lines;
|
||||||
|
const vtt = parseVtt(text);
|
||||||
|
if (vtt.length) return vtt;
|
||||||
|
return parseXmlCaptions(text);
|
||||||
|
}
|
||||||
|
const vtt = parseVtt(text);
|
||||||
|
if (vtt.length) return vtt;
|
||||||
|
const j = asJson3();
|
||||||
|
if (j.length) return j;
|
||||||
|
return parseXmlCaptions(text);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Map short (`yt`) and long (`youtube`) provider ids to `providerUrlFrom()` names. */
|
||||||
|
export function normalizeTranscriptProvider(provider) {
|
||||||
|
const p = String(provider || '').trim().toLowerCase();
|
||||||
|
const map = {
|
||||||
|
yt: 'youtube', youtube: 'youtube',
|
||||||
|
dm: 'dailymotion', dailymotion: 'dailymotion',
|
||||||
|
tw: 'twitch', twitch: 'twitch',
|
||||||
|
pt: 'peertube', peertube: 'peertube',
|
||||||
|
od: 'odysee', odysee: 'odysee',
|
||||||
|
ru: 'rumble', rumble: 'rumble',
|
||||||
|
};
|
||||||
|
return map[p] || null;
|
||||||
|
}
|
||||||
|
|
||||||
|
export { trackExtOf as transcriptTrackExt, firstTrack as transcriptFirstTrack };
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
import { Injectable, inject } from '@angular/core';
|
||||||
|
import { HttpClient } from '@angular/common/http';
|
||||||
|
import { Observable, of } from 'rxjs';
|
||||||
|
import { tap } from 'rxjs/operators';
|
||||||
|
import { dedupeSort } from './suggest.util';
|
||||||
|
|
||||||
|
export interface SuggestResponse {
|
||||||
|
q: string;
|
||||||
|
groups: Record<string, string[]>;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface SuggestCacheEntry {
|
||||||
|
t: number;
|
||||||
|
data: SuggestResponse;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Step 15 — Query typeahead service.
|
||||||
|
*
|
||||||
|
* Contract: no network while `q.trim().length < 2`. Callers debounce
|
||||||
|
* (250-300ms) + `switchMap` (in-flight cancellation). Responses are cached
|
||||||
|
* 5 minutes, deduped and sorted per provider group.
|
||||||
|
*/
|
||||||
|
@Injectable({ providedIn: 'root' })
|
||||||
|
export class SuggestService {
|
||||||
|
private http = inject(HttpClient);
|
||||||
|
private cache = new Map<string, SuggestCacheEntry>();
|
||||||
|
private cacheTtlMs = 5 * 60 * 1000;
|
||||||
|
|
||||||
|
private apiBase(): string {
|
||||||
|
try {
|
||||||
|
const port = window?.location?.port || '';
|
||||||
|
return port && port !== '4000' ? '/proxy/api' : '/api';
|
||||||
|
} catch {
|
||||||
|
return '/api';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Single suggest round-trip (no debounce here — the caller owns it). */
|
||||||
|
suggestOnce(q: string, providers?: string[] | 'all', limit = 10): Observable<SuggestResponse> {
|
||||||
|
const query = String(q ?? '').trim();
|
||||||
|
if (query.length < 2) return of({ q: query, groups: {} });
|
||||||
|
const prov = Array.isArray(providers) && providers.length > 0 ? [...providers].sort().join(',') : 'all';
|
||||||
|
const key = `${prov}|${query.toLowerCase()}|${limit}`;
|
||||||
|
const now = Date.now();
|
||||||
|
const hit = this.cache.get(key);
|
||||||
|
if (hit && now - hit.t < this.cacheTtlMs) return of(hit.data);
|
||||||
|
|
||||||
|
const params = new URLSearchParams({ q: query, limit: String(limit) });
|
||||||
|
if (Array.isArray(providers) && providers.length > 0) params.set('providers', providers.join(','));
|
||||||
|
return this.http.get<SuggestResponse>(`${this.apiBase()}/search/suggest?${params.toString()}`).pipe(
|
||||||
|
tap((res) => {
|
||||||
|
try {
|
||||||
|
const groups: Record<string, string[]> = {};
|
||||||
|
for (const [pid, list] of Object.entries(res?.groups || {})) {
|
||||||
|
groups[pid] = dedupeSort(Array.isArray(list) ? list : [], query).slice(0, limit);
|
||||||
|
}
|
||||||
|
const clean: SuggestResponse = { q: res?.q ?? query, groups };
|
||||||
|
if (this.cache.size > 200) this.cache.clear();
|
||||||
|
this.cache.set(key, { t: Date.now(), data: clean });
|
||||||
|
// Mutate in place so subscribers see the cleaned payload
|
||||||
|
res.q = clean.q;
|
||||||
|
res.groups = clean.groups;
|
||||||
|
} catch {}
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
clearCache(): void {
|
||||||
|
this.cache.clear();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
import 'zone.js/node';
|
||||||
|
import '@angular/compiler'; // JIT needed: decorators are evaluated at import time under ts-node
|
||||||
|
import { firstValueFrom, of } from 'rxjs';
|
||||||
|
import { Injector, runInInjectionContext } from '@angular/core';
|
||||||
|
import { HttpClient } from '@angular/common/http';
|
||||||
|
import { SuggestService } from './suggest.service';
|
||||||
|
import { dedupeSort, mergeGroups, highlightParts } from './suggest.util';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Step 15 — Unit tests for query typeahead (parsing / dedup / min-length guard).
|
||||||
|
*
|
||||||
|
* Same harness as Step 11: ts-node isolated transpilation, manual injection
|
||||||
|
* with a minimal HttpClient fake returning real Observables.
|
||||||
|
* Run with: npm run test:suggest
|
||||||
|
*/
|
||||||
|
|
||||||
|
function assertEqual(actual: unknown, expected: unknown, message: string): void {
|
||||||
|
const a = JSON.stringify(actual);
|
||||||
|
const b = JSON.stringify(expected);
|
||||||
|
if (a !== b) {
|
||||||
|
throw new Error(`${message} (expected ${b}, got ${a})`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function assert(condition: unknown, message: string): void {
|
||||||
|
if (!condition) {
|
||||||
|
throw new Error(message);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function logOk(msg: string): void {
|
||||||
|
console.log(`✓ ${msg}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
(async () => {
|
||||||
|
// --- Pure helpers: dedupe + sort ---
|
||||||
|
assertEqual(
|
||||||
|
dedupeSort(['Tutoriel Angular', 'tutoriel angular ', ' ', 'Tutoriel Android'], 'tutoriel an'),
|
||||||
|
['Tutoriel Android', 'Tutoriel Angular'],
|
||||||
|
'dedupeSort dedupes case-insensitively and sorts prefix-first then alpha',
|
||||||
|
);
|
||||||
|
logOk('dedupeSort dedup + tri');
|
||||||
|
|
||||||
|
assertEqual(dedupeSort(['b', 'a', 'c']), ['a', 'b', 'c'], 'dedupeSort sorts alphabetically without query');
|
||||||
|
logOk('dedupeSort alpha fallback');
|
||||||
|
|
||||||
|
// --- Pure helpers: merge groups ---
|
||||||
|
const merged = mergeGroups({ dm: ['Tutoriel android'], yt: ['Tutoriel Angular', 'tutoriel android'], xx: ['zz'] });
|
||||||
|
assertEqual(
|
||||||
|
merged,
|
||||||
|
[
|
||||||
|
{ text: 'Tutoriel Angular', provider: 'yt' },
|
||||||
|
{ text: 'tutoriel android', provider: 'yt' },
|
||||||
|
{ text: 'zz', provider: 'xx' },
|
||||||
|
],
|
||||||
|
'mergeGroups respects registry order, dedupes across providers, unknowns last',
|
||||||
|
);
|
||||||
|
logOk('mergeGroups order + dedup');
|
||||||
|
|
||||||
|
// --- Pure helpers: highlight ---
|
||||||
|
assertEqual(
|
||||||
|
highlightParts('Tutoriel Angular', 'toriel an'),
|
||||||
|
{ pre: 'Tu', match: 'toriel An', post: 'gular' },
|
||||||
|
'highlightParts splits on first case-insensitive occurrence',
|
||||||
|
);
|
||||||
|
assertEqual(
|
||||||
|
highlightParts('Bonjour', 'zzz'),
|
||||||
|
{ pre: 'Bonjour', match: '', post: '' },
|
||||||
|
'highlightParts without occurrence returns plain text',
|
||||||
|
);
|
||||||
|
logOk('highlightParts');
|
||||||
|
|
||||||
|
// --- Service: min length guard (no network below 2 chars) ---
|
||||||
|
let httpCalls = 0;
|
||||||
|
const httpFake = {
|
||||||
|
get: (url: string) => {
|
||||||
|
httpCalls += 1;
|
||||||
|
assert(url.includes('/search/suggest'), `suggest hits /api/search/suggest (got ${url})`);
|
||||||
|
return of({ q: 'angular', groups: { yt: ['angular tutorial', 'Angular'], dm: ['angular tutorial'] } });
|
||||||
|
},
|
||||||
|
};
|
||||||
|
const injector = Injector.create([{ provide: HttpClient, useValue: httpFake }]);
|
||||||
|
const svc: SuggestService = runInInjectionContext(injector, () => new SuggestService());
|
||||||
|
|
||||||
|
const short = await firstValueFrom(svc.suggestOnce('a'));
|
||||||
|
assertEqual(short, { q: 'a', groups: {} }, 'q < 2 chars returns empty without network');
|
||||||
|
assertEqual(httpCalls, 0, 'no HTTP call for short queries');
|
||||||
|
logOk('suggestOnce min-length guard (no network)');
|
||||||
|
|
||||||
|
// --- Service: dedup + cache ---
|
||||||
|
const first = await firstValueFrom(svc.suggestOnce('angular', ['yt', 'dm']));
|
||||||
|
assertEqual(httpCalls, 1, 'one HTTP call for a fresh query');
|
||||||
|
assertEqual(first.groups['yt'], ['Angular', 'angular tutorial'], 'per-provider groups deduped + sorted');
|
||||||
|
const second = await firstValueFrom(svc.suggestOnce('angular', ['dm', 'yt']));
|
||||||
|
assertEqual(httpCalls, 1, 'second identical query served from cache (provider order-insensitive)');
|
||||||
|
assert(second === first || JSON.stringify(second) === JSON.stringify(first), 'cached payload matches');
|
||||||
|
logOk('suggestOnce dedup + cache 5 min');
|
||||||
|
|
||||||
|
console.log('\nAll suggest tests passed.');
|
||||||
|
})().catch((e) => {
|
||||||
|
console.error(e instanceof Error ? e.stack : String(e));
|
||||||
|
process.exit(1);
|
||||||
|
});
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
/**
|
||||||
|
* Step 15 — Pure helpers for query typeahead (testable without Angular).
|
||||||
|
*/
|
||||||
|
|
||||||
|
/** Normalize a suggestion string for dedup comparisons. */
|
||||||
|
export function normSuggest(s: string): string {
|
||||||
|
return String(s ?? '').trim().toLowerCase().replace(/\s+/g, ' ');
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Deduplicate (case-insensitive, first occurrence wins) then sort:
|
||||||
|
* items starting with `q` come first, alphabetical within each group.
|
||||||
|
*/
|
||||||
|
export function dedupeSort(items: string[], q = ''): string[] {
|
||||||
|
const seen = new Set<string>();
|
||||||
|
const clean: string[] = [];
|
||||||
|
for (const raw of items || []) {
|
||||||
|
const t = String(raw ?? '').trim();
|
||||||
|
if (!t) continue;
|
||||||
|
const k = normSuggest(t);
|
||||||
|
if (seen.has(k)) continue;
|
||||||
|
seen.add(k);
|
||||||
|
clean.push(t);
|
||||||
|
}
|
||||||
|
const nq = normSuggest(q);
|
||||||
|
clean.sort((a, b) => {
|
||||||
|
const na = normSuggest(a);
|
||||||
|
const nb = normSuggest(b);
|
||||||
|
const pa = nq ? (na.startsWith(nq) ? 0 : 1) : 0;
|
||||||
|
const pb = nq ? (nb.startsWith(nq) ? 0 : 1) : 0;
|
||||||
|
if (pa !== pb) return pa - pb;
|
||||||
|
return na < nb ? -1 : na > nb ? 1 : 0;
|
||||||
|
});
|
||||||
|
return clean;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Flatten provider groups into a single ordered list (registry order),
|
||||||
|
* deduped case-insensitively. Unknown provider keys go last (insertion order).
|
||||||
|
*/
|
||||||
|
export function mergeGroups(
|
||||||
|
groups: Record<string, string[]> | null | undefined,
|
||||||
|
order: string[] = ['yt', 'dm', 'tw', 'pt', 'od', 'ru'],
|
||||||
|
): Array<{ text: string; provider: string }> {
|
||||||
|
const out: Array<{ text: string; provider: string }> = [];
|
||||||
|
const seen = new Set<string>();
|
||||||
|
const push = (pid: string, list: string[] | undefined) => {
|
||||||
|
for (const raw of list || []) {
|
||||||
|
const t = String(raw ?? '').trim();
|
||||||
|
if (!t) continue;
|
||||||
|
const k = normSuggest(t);
|
||||||
|
if (seen.has(k)) continue;
|
||||||
|
seen.add(k);
|
||||||
|
out.push({ text: t, provider: pid });
|
||||||
|
}
|
||||||
|
};
|
||||||
|
for (const pid of order) push(pid, groups?.[pid]);
|
||||||
|
for (const pid of Object.keys(groups || {})) {
|
||||||
|
if (!order.includes(pid)) push(pid, (groups as Record<string, string[]>)[pid]);
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Split `text` on the first case-insensitive occurrence of `q` for <mark> highlighting. */
|
||||||
|
export function highlightParts(text: string, q: string): { pre: string; match: string; post: string } {
|
||||||
|
const t = String(text ?? '');
|
||||||
|
const needle = String(q ?? '').trim();
|
||||||
|
if (!needle) return { pre: t, match: '', post: '' };
|
||||||
|
const idx = t.toLowerCase().indexOf(needle.toLowerCase());
|
||||||
|
if (idx === -1) return { pre: t, match: '', post: '' };
|
||||||
|
return { pre: t.slice(0, idx), match: t.slice(idx, idx + needle.length), post: t.slice(idx + needle.length) };
|
||||||
|
}
|
||||||
@@ -23,6 +23,7 @@ import { formatAbsoluteFr, formatNumberFr } from '../../utils/date.util';
|
|||||||
import { ProviderBadgeComponent } from '../../app/shared/components/provider-badge/provider-badge.component';
|
import { ProviderBadgeComponent } from '../../app/shared/components/provider-badge/provider-badge.component';
|
||||||
import { ChannelIdentityComponent } from '../../app/shared/components/channel-identity/channel-identity.component';
|
import { ChannelIdentityComponent } from '../../app/shared/components/channel-identity/channel-identity.component';
|
||||||
import { DurationPipe } from '../../app/shared/pipes/duration.pipe';
|
import { DurationPipe } from '../../app/shared/pipes/duration.pipe';
|
||||||
|
import { UserService } from '../../services/user.service';
|
||||||
|
|
||||||
@Component({
|
@Component({
|
||||||
selector: 'app-watch',
|
selector: 'app-watch',
|
||||||
@@ -44,6 +45,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
private auth = inject(AuthService);
|
private auth = inject(AuthService);
|
||||||
private http = inject(HttpClient);
|
private http = inject(HttpClient);
|
||||||
private iframeProgress = inject(IframeProgressService);
|
private iframeProgress = inject(IframeProgressService);
|
||||||
|
private users = inject(UserService);
|
||||||
private channels = inject(ChannelsService);
|
private channels = inject(ChannelsService);
|
||||||
private subs = inject(SubscriptionsService);
|
private subs = inject(SubscriptionsService);
|
||||||
private routeSubscription!: Subscription;
|
private routeSubscription!: Subscription;
|
||||||
@@ -197,6 +199,8 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
// Build a direct Twitch URL as a fallback open-in-new-tab action
|
// Build a direct Twitch URL as a fallback open-in-new-tab action
|
||||||
twitchOpenUrl(): string | null {
|
twitchOpenUrl(): string | null {
|
||||||
if (this.provider() !== 'twitch') return null;
|
if (this.provider() !== 'twitch') return null;
|
||||||
|
const clip = this.twitchClip();
|
||||||
|
if (clip) return `https://clips.twitch.tv/${encodeURIComponent(clip)}`;
|
||||||
const ch = this.twitchChannel();
|
const ch = this.twitchChannel();
|
||||||
const id = this.videoId();
|
const id = this.videoId();
|
||||||
const cur = this.video();
|
const cur = this.video();
|
||||||
@@ -286,6 +290,20 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
liked = signal<boolean>(false);
|
liked = signal<boolean>(false);
|
||||||
likeBusy = signal<boolean>(false);
|
likeBusy = signal<boolean>(false);
|
||||||
|
|
||||||
|
// --- Transcript state (Step 16, Phase 1: display + language selector, no seek) ---
|
||||||
|
transcriptOpen = signal(false);
|
||||||
|
transcriptLoading = signal(false);
|
||||||
|
transcriptAvailable = signal(false);
|
||||||
|
transcriptError = signal<string | null>(null);
|
||||||
|
// Machine-readable unavailability reason from the API: 'no_subtitles' |
|
||||||
|
// 'provider_unsupported' | 'temporarily_unavailable' | null.
|
||||||
|
transcriptReason = signal<string | null>(null);
|
||||||
|
// True when retrying may succeed (transient YouTube rate-limit/fetch failure).
|
||||||
|
transcriptRetryable = signal(false);
|
||||||
|
transcriptLanguages = signal<string[]>([]);
|
||||||
|
transcriptLines = signal<Array<{ t: number; dur: number; text: string }>>([]);
|
||||||
|
selectedTranscriptLang = signal('fr');
|
||||||
|
|
||||||
// --- Watch history tracking for native player ---
|
// --- Watch history tracking for native player ---
|
||||||
watchHistoryId = signal<string | null>(null);
|
watchHistoryId = signal<string | null>(null);
|
||||||
startPositionSeconds = signal<number | null>(null);
|
startPositionSeconds = signal<number | null>(null);
|
||||||
@@ -296,6 +314,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
private providerSel = signal<string>('');
|
private providerSel = signal<string>('');
|
||||||
provider = computed(() => this.providerSel() || this.instances.selectedProvider());
|
provider = computed(() => this.providerSel() || this.instances.selectedProvider());
|
||||||
private twitchChannel = signal<string | null>(null);
|
private twitchChannel = signal<string | null>(null);
|
||||||
|
private twitchClip = signal<string | null>(null);
|
||||||
// Backend-provided embed URL for providers that need it (e.g., Rumble)
|
// Backend-provided embed URL for providers that need it (e.g., Rumble)
|
||||||
private rumbleEmbed = signal<string | null>(null);
|
private rumbleEmbed = signal<string | null>(null);
|
||||||
|
|
||||||
@@ -303,9 +322,10 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
const id = this.videoId();
|
const id = this.videoId();
|
||||||
const p = this.provider();
|
const p = this.provider();
|
||||||
const ch = this.twitchChannel();
|
const ch = this.twitchChannel();
|
||||||
|
const clip = this.twitchClip();
|
||||||
const slug = this.odyseeSlug();
|
const slug = this.odyseeSlug();
|
||||||
const start = Math.max(0, this.startPositionSeconds() || 0);
|
const start = Math.max(0, this.startPositionSeconds() || 0);
|
||||||
if (!id && !(p === 'twitch' && ch) && !(p === 'odysee' && slug)) return null;
|
if (!id && !(p === 'twitch' && (ch || clip)) && !(p === 'odysee' && slug)) return null;
|
||||||
try {
|
try {
|
||||||
const host = (location && location.hostname) ? location.hostname : 'localhost';
|
const host = (location && location.hostname) ? location.hostname : 'localhost';
|
||||||
if (p === 'youtube') {
|
if (p === 'youtube') {
|
||||||
@@ -324,9 +344,15 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
return this.sanitizer.bypassSecurityTrustResourceUrl(u);
|
return this.sanitizer.bypassSecurityTrustResourceUrl(u);
|
||||||
}
|
}
|
||||||
if (p === 'twitch') {
|
if (p === 'twitch') {
|
||||||
// Channel or VOD; Twitch requires one or more 'parent' params that match the embedding domain
|
// Channel, clip ou VOD ; Twitch exige 'parent' = domaine d'embed.
|
||||||
const parentsSet = new Set<string>([host, 'localhost', '127.0.0.1']);
|
const parentsSet = new Set<string>([host, 'localhost', '127.0.0.1']);
|
||||||
const parentParams = Array.from(parentsSet).map(h => `parent=${encodeURIComponent(h)}`).join('&');
|
const parentParams = Array.from(parentsSet).map(h => `parent=${encodeURIComponent(h)}`).join('&');
|
||||||
|
// Clips : player dédié clips.twitch.tv (le player VOD rejette les slugs).
|
||||||
|
const clipId = clip || (/^\d+$/.test(String(id || '')) ? null : ((this.video() as any)?.kind === 'clip' ? id : null));
|
||||||
|
if (clipId) {
|
||||||
|
const u = `https://clips.twitch.tv/embed?clip=${encodeURIComponent(clipId)}&${parentParams}&autoplay=false`;
|
||||||
|
return this.sanitizer.bypassSecurityTrustResourceUrl(u);
|
||||||
|
}
|
||||||
const base = 'https://player.twitch.tv/';
|
const base = 'https://player.twitch.tv/';
|
||||||
// ?video= exige un id de VOD numérique ; un login (non numérique) sans
|
// ?video= exige un id de VOD numérique ; un login (non numérique) sans
|
||||||
// ?channel= doit être traité comme une chaîne (résultats de recherche).
|
// ?channel= doit être traité comme une chaîne (résultats de recherche).
|
||||||
@@ -461,11 +487,13 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
const slug = this.route.snapshot.queryParamMap.get('slug');
|
const slug = this.route.snapshot.queryParamMap.get('slug');
|
||||||
const p = (this.route.snapshot.queryParamMap.get('p') || providerFromPath) as Provider | null;
|
const p = (this.route.snapshot.queryParamMap.get('p') || providerFromPath) as Provider | null;
|
||||||
const ch = this.route.snapshot.queryParamMap.get('channel');
|
const ch = this.route.snapshot.queryParamMap.get('channel');
|
||||||
|
const clip = this.route.snapshot.queryParamMap.get('clip');
|
||||||
if (id) {
|
if (id) {
|
||||||
this.videoId.set(id);
|
this.videoId.set(id);
|
||||||
this.odyseeSlug.set(slug);
|
this.odyseeSlug.set(slug);
|
||||||
this.providerSel.set(p || '');
|
this.providerSel.set(p || '');
|
||||||
this.twitchChannel.set(ch);
|
this.twitchChannel.set(ch);
|
||||||
|
this.twitchClip.set(clip);
|
||||||
|
|
||||||
// Choisir un provider final (paramètre présent, segment d'URL, ou fallback depuis état courant)
|
// Choisir un provider final (paramètre présent, segment d'URL, ou fallback depuis état courant)
|
||||||
let finalProvider = p as Provider | null;
|
let finalProvider = p as Provider | null;
|
||||||
@@ -583,6 +611,7 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
this.summaryError.set(null);
|
this.summaryError.set(null);
|
||||||
this.selectedQuality.set(null);
|
this.selectedQuality.set(null);
|
||||||
this.resetDownloadUi();
|
this.resetDownloadUi();
|
||||||
|
this.resetTranscriptUi();
|
||||||
// Reset like state while loading
|
// Reset like state while loading
|
||||||
this.liked.set(false);
|
this.liked.set(false);
|
||||||
this.likeBusy.set(false);
|
this.likeBusy.set(false);
|
||||||
@@ -861,6 +890,189 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
this.selectedQuality.set(value || null);
|
this.selectedQuality.set(value || null);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- Transcript logic (Step 16) ---
|
||||||
|
private resetTranscriptUi() {
|
||||||
|
this.transcriptOpen.set(false);
|
||||||
|
this.transcriptLoading.set(false);
|
||||||
|
this.transcriptAvailable.set(false);
|
||||||
|
this.transcriptError.set(null);
|
||||||
|
this.transcriptReason.set(null);
|
||||||
|
this.transcriptRetryable.set(false);
|
||||||
|
this.transcriptLanguages.set([]);
|
||||||
|
this.transcriptLines.set([]);
|
||||||
|
try {
|
||||||
|
const prefLang = String(this.users.preferences()?.language || '').trim().toLowerCase().slice(0, 12);
|
||||||
|
this.selectedTranscriptLang.set(prefLang || 'fr');
|
||||||
|
} catch {
|
||||||
|
this.selectedTranscriptLang.set('fr');
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
toggleTranscript() {
|
||||||
|
const next = !this.transcriptOpen();
|
||||||
|
this.transcriptOpen.set(next);
|
||||||
|
if (next && this.transcriptLines().length === 0 && !this.transcriptLoading()) {
|
||||||
|
this.loadTranscript();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
onTranscriptLangChange(value: string) {
|
||||||
|
const lang = String(value || '').trim().slice(0, 12) || 'fr';
|
||||||
|
this.selectedTranscriptLang.set(lang);
|
||||||
|
this.loadTranscript();
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Map a backend transcript state (reason/error/retryable) to the UI:
|
||||||
|
* user-facing message + whether a "Réessayer" button makes sense. */
|
||||||
|
private transcriptStatusFor(data: any): { message: string | null; retryable: boolean; reason: string | null } {
|
||||||
|
const reason = typeof data?.reason === 'string' ? data.reason : null;
|
||||||
|
const code = typeof data?.error === 'string' ? data.error : '';
|
||||||
|
const retryable = data?.retryable === true
|
||||||
|
|| code === 'transcript_temporarily_unavailable'
|
||||||
|
|| code === 'rate_limited';
|
||||||
|
if (reason === 'provider_unsupported') {
|
||||||
|
return {
|
||||||
|
message: 'Les transcripts ne sont pas pris en charge pour ce fournisseur (aucun sous-titre exposé).',
|
||||||
|
retryable: false,
|
||||||
|
reason,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
if (reason === 'temporarily_unavailable' || code === 'transcript_temporarily_unavailable') {
|
||||||
|
return {
|
||||||
|
message: 'Sous-titres temporairement indisponibles (limite YouTube). Réessayez dans quelques instants.',
|
||||||
|
retryable: true,
|
||||||
|
reason: reason || 'temporarily_unavailable',
|
||||||
|
};
|
||||||
|
}
|
||||||
|
if (reason === 'no_subtitles') {
|
||||||
|
return { message: null, retryable: false, reason };
|
||||||
|
}
|
||||||
|
// Legacy / unknown error code: sanitize, never leak HTML or parse noise.
|
||||||
|
const clean = this.transcriptDisplayError(code);
|
||||||
|
if (clean) return { message: `Transcript indisponible. (${clean})`, retryable, reason };
|
||||||
|
if (retryable) {
|
||||||
|
return {
|
||||||
|
message: 'Sous-titres temporairement indisponibles. Réessayez dans quelques instants.',
|
||||||
|
retryable: true,
|
||||||
|
reason: reason || 'temporarily_unavailable',
|
||||||
|
};
|
||||||
|
}
|
||||||
|
return { message: null, retryable: false, reason };
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Sanitize a backend error code for display: never leak HTML pages or raw
|
||||||
|
* JSON-parse noise (e.g. proxy / CDN "<!DOCTYPE..." bodies surfaced as
|
||||||
|
* "SyntaxError: Unexpected token '<'..."). Returns null when nothing
|
||||||
|
* user-friendly can be shown (the "no subtitles" notice already covers it). */
|
||||||
|
private transcriptDisplayError(raw: unknown): string | null {
|
||||||
|
let msg = '';
|
||||||
|
try {
|
||||||
|
if (typeof raw === 'string') msg = raw;
|
||||||
|
else if (raw instanceof Error) msg = raw.message || '';
|
||||||
|
else if (raw != null) msg = String(raw);
|
||||||
|
} catch {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
msg = msg.trim().slice(0, 160);
|
||||||
|
if (!msg) return null;
|
||||||
|
const lowered = msg.toLowerCase();
|
||||||
|
// Ignore internal codes handled silently + any HTML / JSON-parse noise.
|
||||||
|
if (
|
||||||
|
msg === 'transcript_fetch_failed' ||
|
||||||
|
msg === 'transcript_temporarily_unavailable' ||
|
||||||
|
msg === 'rate_limited' ||
|
||||||
|
lowered.includes('<!doctype') ||
|
||||||
|
lowered.includes('<html') ||
|
||||||
|
msg.includes('<') ||
|
||||||
|
lowered.includes('not valid json') ||
|
||||||
|
lowered.includes('unexpected token') ||
|
||||||
|
lowered.startsWith('syntaxerror')
|
||||||
|
) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
return msg;
|
||||||
|
}
|
||||||
|
|
||||||
|
retryTranscript() {
|
||||||
|
if (this.transcriptLoading()) return;
|
||||||
|
this.transcriptLines.set([]);
|
||||||
|
this.loadTranscript();
|
||||||
|
}
|
||||||
|
|
||||||
|
loadTranscript() {
|
||||||
|
const p = this.provider();
|
||||||
|
const id = this.videoId();
|
||||||
|
if (!p || !id) {
|
||||||
|
this.transcriptError.set('Vidéo introuvable.');
|
||||||
|
this.transcriptReason.set(null);
|
||||||
|
this.transcriptRetryable.set(false);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
// Never break the Watch page: every failure degrades to the "unavailable" notice.
|
||||||
|
try {
|
||||||
|
this.transcriptLoading.set(true);
|
||||||
|
this.transcriptError.set(null);
|
||||||
|
this.transcriptReason.set(null);
|
||||||
|
this.transcriptRetryable.set(false);
|
||||||
|
const opts = this.buildProviderOpts() as Record<string, string>;
|
||||||
|
const params = new URLSearchParams({ lang: this.selectedTranscriptLang() });
|
||||||
|
if (opts['instance']) params.set('instance', opts['instance']);
|
||||||
|
if (opts['slug']) params.set('slug', opts['slug']);
|
||||||
|
const url = `${this.apiBase()}/transcript/${encodeURIComponent(p)}/${encodeURIComponent(id)}?${params.toString()}`;
|
||||||
|
this.http.get<any>(url).subscribe({
|
||||||
|
next: (data) => {
|
||||||
|
this.transcriptLoading.set(false);
|
||||||
|
this.transcriptAvailable.set(!!data?.available);
|
||||||
|
this.transcriptLanguages.set(Array.isArray(data?.languages) ? data.languages : []);
|
||||||
|
this.transcriptLines.set(Array.isArray(data?.lines) ? data.lines : []);
|
||||||
|
if (data?.available && typeof data?.lang === 'string' && data.lang) {
|
||||||
|
this.selectedTranscriptLang.set(data.lang);
|
||||||
|
}
|
||||||
|
if (!data?.available) {
|
||||||
|
const st = this.transcriptStatusFor(data);
|
||||||
|
this.transcriptReason.set(st.reason);
|
||||||
|
this.transcriptRetryable.set(st.retryable);
|
||||||
|
this.transcriptError.set(st.message);
|
||||||
|
} else {
|
||||||
|
this.transcriptReason.set(null);
|
||||||
|
this.transcriptRetryable.set(false);
|
||||||
|
this.transcriptError.set(null);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
error: (err) => {
|
||||||
|
this.transcriptLoading.set(false);
|
||||||
|
this.transcriptAvailable.set(false);
|
||||||
|
this.transcriptLines.set([]);
|
||||||
|
const body = err?.error;
|
||||||
|
if (body && typeof body === 'object' && Array.isArray(body.languages)) {
|
||||||
|
this.transcriptLanguages.set(body.languages);
|
||||||
|
}
|
||||||
|
const raw =
|
||||||
|
(body && typeof body === 'object' && (body.error || body.details || body.reason)) ||
|
||||||
|
(typeof body === 'string' ? body : '') ||
|
||||||
|
err?.message ||
|
||||||
|
'';
|
||||||
|
const status = err?.status;
|
||||||
|
if (status === 429 || String(raw || '').trim() === 'rate_limited') {
|
||||||
|
this.transcriptReason.set('temporarily_unavailable');
|
||||||
|
this.transcriptRetryable.set(true);
|
||||||
|
this.transcriptError.set('Trop de requêtes. Réessayez dans une minute.');
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const st = this.transcriptStatusFor(
|
||||||
|
body && typeof body === 'object' ? body : { error: raw },
|
||||||
|
);
|
||||||
|
this.transcriptReason.set(st.reason);
|
||||||
|
this.transcriptRetryable.set(st.retryable);
|
||||||
|
this.transcriptError.set(st.message);
|
||||||
|
},
|
||||||
|
});
|
||||||
|
} catch {
|
||||||
|
this.transcriptLoading.set(false);
|
||||||
|
this.transcriptAvailable.set(false);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// --- Download logic ---
|
// --- Download logic ---
|
||||||
openDownloadPanel() {
|
openDownloadPanel() {
|
||||||
this.downloadOpen.set(true);
|
this.downloadOpen.set(true);
|
||||||
@@ -1053,6 +1265,10 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
if (slug.startsWith('/')) slug = slug.slice(1);
|
if (slug.startsWith('/')) slug = slug.slice(1);
|
||||||
qp.slug = slug;
|
qp.slug = slug;
|
||||||
}
|
}
|
||||||
|
if (p === 'twitch' && String((v as any).kind || '').toLowerCase() === 'clip') {
|
||||||
|
qp.clip = v.videoId;
|
||||||
|
return qp;
|
||||||
|
}
|
||||||
if (p === 'twitch' && (v.type === 'channel' || (v as any).type === 'live')) {
|
if (p === 'twitch' && (v.type === 'channel' || (v as any).type === 'live')) {
|
||||||
// Prefer explicit videoId as channel login; fallback to parsing URL
|
// Prefer explicit videoId as channel login; fallback to parsing URL
|
||||||
if (v.videoId && !/^\d+$/.test(String(v.videoId))) {
|
if (v.videoId && !/^\d+$/.test(String(v.videoId))) {
|
||||||
@@ -1072,18 +1288,64 @@ export class WatchComponent implements OnDestroy, AfterViewInit {
|
|||||||
private loadRelatedSuggestions(): void {
|
private loadRelatedSuggestions(): void {
|
||||||
const cur = this.video();
|
const cur = this.video();
|
||||||
if (!cur) return;
|
if (!cur) return;
|
||||||
const title = (cur.title || '').trim();
|
|
||||||
const maxItems = 12;
|
const maxItems = 12;
|
||||||
if (title) {
|
const apply = (items: any[]) => {
|
||||||
this.apiService.searchVideosPage(title).subscribe(res => {
|
// Ignore les réponses arrivées après une navigation (race).
|
||||||
const items = (res.items || []).filter(v => v.videoId !== cur.videoId).slice(0, maxItems);
|
try {
|
||||||
this.video.update(v => v ? { ...v, relatedStreams: items } : v);
|
if ((this.videoId() || '') !== (cur.videoId || '')) return;
|
||||||
});
|
} catch {}
|
||||||
} else {
|
const list = (items || []).filter((v: any) => (v?.videoId || v?.id) !== cur.videoId).slice(0, maxItems);
|
||||||
this.apiService.getTrendingPage().subscribe(res => {
|
this.video.update(v => v ? { ...v, relatedStreams: list } : v);
|
||||||
const items = (res.items || []).filter(v => v.videoId !== cur.videoId).slice(0, maxItems);
|
};
|
||||||
this.video.update(v => v ? { ...v, relatedStreams: items } : v);
|
const titleFallback = () => {
|
||||||
});
|
const title = (cur.title || '').trim();
|
||||||
}
|
if (title) {
|
||||||
|
this.apiService.searchVideosPage(title).subscribe(res => {
|
||||||
|
apply(res.items || []);
|
||||||
|
});
|
||||||
|
} else {
|
||||||
|
this.apiService.getTrendingPage().subscribe(res => {
|
||||||
|
apply(res.items || []);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
};
|
||||||
|
// YouTube : vrai watch-next InnerTube via /api/details (0 quota, 0 recherche
|
||||||
|
// supplémentaire) — façon SmartTube. Autres providers : fallback existant.
|
||||||
|
try {
|
||||||
|
if (this.provider() === 'youtube' && cur.videoId) {
|
||||||
|
const id = cur.videoId;
|
||||||
|
this.http.get<any>(`/api/details/youtube/${encodeURIComponent(id)}`).subscribe({
|
||||||
|
next: (data: any) => {
|
||||||
|
const rel = Array.isArray(data?.related) ? data.related : [];
|
||||||
|
if (rel.length > 0) {
|
||||||
|
apply(rel.map((it: any) => ({
|
||||||
|
url: it.url || `https://www.youtube.com/watch?v=${it.id}`,
|
||||||
|
type: 'video',
|
||||||
|
title: it.title || '',
|
||||||
|
thumbnail: it.thumbnail || '',
|
||||||
|
uploaderName: it.uploaderName || '',
|
||||||
|
uploaderAvatar: '',
|
||||||
|
channelId: it.channelId || undefined,
|
||||||
|
channelHandle: it.channelHandle || undefined,
|
||||||
|
channelExternalId: it.channelExternalId || it.channelId || undefined,
|
||||||
|
uploadedDate: it.publishedAt || '',
|
||||||
|
duration: typeof it.duration === 'number' ? it.duration : 0,
|
||||||
|
views: typeof it.views === 'number' ? it.views : 0,
|
||||||
|
uploaded: 0,
|
||||||
|
videoId: String(it.id || ''),
|
||||||
|
provider: 'youtube',
|
||||||
|
isShort: it.isShort === true ? true : undefined,
|
||||||
|
isLive: it.isLive === true ? true : undefined,
|
||||||
|
} as Video)));
|
||||||
|
} else {
|
||||||
|
titleFallback();
|
||||||
|
}
|
||||||
|
},
|
||||||
|
error: () => titleFallback(),
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
} catch {}
|
||||||
|
titleFallback();
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -17,6 +17,179 @@
|
|||||||
- [x] Step 12: Write basic e2e scenarios for providers filtering and deep-link
|
- [x] Step 12: Write basic e2e scenarios for providers filtering and deep-link
|
||||||
- [x] Step 13: Add minimal telemetry hooks and events
|
- [x] Step 13: Add minimal telemetry hooks and events
|
||||||
- [x] Step 14: Update README with search UX section, keyboard shortcuts and usage notes
|
- [x] Step 14: Update README with search UX section, keyboard shortcuts and usage notes
|
||||||
|
- [x] Step 15: Auto-complétion de la **requête** (typeahead) dans la barre de recherche
|
||||||
|
- [x] Step 16: **Transcripts** de vidéos sur la page Watch (API + UI, Phase 1 du doc)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Step 15 — Suggestions de requêtes dans la barre de recherche
|
||||||
|
|
||||||
|
### Ce que ça veut dire (et ce que ce n'est pas)
|
||||||
|
|
||||||
|
Quand l'utilisateur tape `tutoriel an`, la barre affiche **avant qu'il valide** une liste de
|
||||||
|
requêtes complètes probables : `tutoriel angular`, `tutoriel android`… + ses recherches passées.
|
||||||
|
|
||||||
|
> ⚠️ À ne pas confondre avec `SearchSuggestionsComponent` (`src/components/search/search-suggestions.component.ts`),
|
||||||
|
> qui lui affiche les **résultats de recherche** groupés par provider **après** la validation sur `/search`.
|
||||||
|
> Le Step 15 ajoute un nouveau panneau de **suggestions de texte** sous l'`<input>`.
|
||||||
|
|
||||||
|
### Ce qui existe déjà (à réutiliser)
|
||||||
|
|
||||||
|
| Élément | Fichier | État |
|
||||||
|
|---|---|---|
|
||||||
|
| `@Output() searchChange` (query debouncée 300 ms) | `src/components/search/search-box.component.ts:46` | Émis mais **aucun abonné** — le header ne branche que `(submitted)` (`header.component.html:19`) |
|
||||||
|
| Popover `@provider` + navigation clavier | `search-box.component.ts:16` (`AT_QUERY_RE`), `.html:56-76` | ✅ mais ne couvre que les ids de providers |
|
||||||
|
| Historique de recherches | `HistoryService.getSearchHistory(n)` | ✅ utilisé dans le Quick Menu |
|
||||||
|
| Cache in-memory TTL 60 s | `src/app/search/search.service.ts:33` | Pattern à dupliquer |
|
||||||
|
| Fan-out multi-providers | `GET /api/search` (`server/index.mjs:2116`) | Pattern à dupliquer |
|
||||||
|
|
||||||
|
### Livrables
|
||||||
|
|
||||||
|
1. **Backend** — `GET /api/search/suggest?q=…&providers=yt,dm&limit=10`
|
||||||
|
→ `{ q, groups: { yt: string[], dm: string[], … } }`
|
||||||
|
- Un `suggest(q, { limit })` optionnel par handler dans `server/providers/*.mjs` + déclaré dans
|
||||||
|
`server/providers/registry.mjs` (`ProviderAdapter`).
|
||||||
|
- Providers sans API de suggestions native → `[]` (dégradation propre, jamais d'erreur 500).
|
||||||
|
- Cache + rate-limit sur le même modèle que `/api/search`.
|
||||||
|
2. **Front service** — `SuggestService` (ou méthode `suggest()` sur `SearchService`) :
|
||||||
|
`min length 2`, `debounce 250-300 ms`, `switchMap` (annulation de la requête précédente),
|
||||||
|
cache 5 min, dédoublonnage + tri.
|
||||||
|
3. **UI** — panneau sous l'input, sections :
|
||||||
|
- 🔎 « Recherches récentes » (History, avec icône horloge)
|
||||||
|
- 📊 Suggestions providers, groupées par provider (badges `YT/DM/…`) ou fusionnées
|
||||||
|
- Sous-chaîne commune mise en évidence
|
||||||
|
4. **Interactions** — clic = remplir + soumettre ; ↑/↓/Enter/Tab/Esc ; **priorité au popover `@`**
|
||||||
|
quand il est ouvert ; `aria-listbox` + `aria-activedescendant` (le `role="combobox"` est déjà sur l'input).
|
||||||
|
5. **Télémétrie** — `suggest_shown`, `suggest_used` (à ajouter à la whitelist serveur de l'étape 13).
|
||||||
|
6. **Tests** — `npm run test:suggest` (unitaires parsing/dédup) + scénario e2e : `providers=yt` ne renvoie
|
||||||
|
que des suggestions `yt`, provider inconnu → fallback registry complet.
|
||||||
|
|
||||||
|
### Critères d'acceptation
|
||||||
|
|
||||||
|
- [x] Aucune requête réseau tant que `q.trim().length < 2`
|
||||||
|
- [x] frappe rapide = une seule requête en vol (pas d'empilement / de réponse qui écrase la dernière)
|
||||||
|
- [x] Esc ferme le panneau sans effacer le texte ; Tab complète sans lancer la recherche
|
||||||
|
- [x] Un provider sans suggestions ou en échec n'affiche ni trou ni erreur dans le panneau
|
||||||
|
- [x] Focus clavier et lecteurs d'écran cohérents avec l'existant (a11y Step 10)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Step 16 — Transcripts (Phase 1 du `docs/transcript_architecture_dev.md`)
|
||||||
|
|
||||||
|
### Principe directeur
|
||||||
|
|
||||||
|
> **Un endpoint, un module de parsing, une UI.** La disponibilité dépend de la plateforme, mais le
|
||||||
|
> **code ne change pas** selon le provider (le doc appelle ça « Rung 2 : réutiliser, pas réinventer »).
|
||||||
|
|
||||||
|
Source des données : `yt-dlp --dump-single-json --skip-download` (déjà utilisé côté serveur, ex.
|
||||||
|
`server/index.mjs:1012`) → on lit `subtitles` / `automatic_captions` → on télécharge la piste `json3` ou `vtt`
|
||||||
|
→ on normalise en `{ lang, available, languages, lines: [{ t, dur, text }] }`.
|
||||||
|
|
||||||
|
### Périmètre Phase 1 (à faire)
|
||||||
|
|
||||||
|
**Backend**
|
||||||
|
|
||||||
|
- [x] `server/transcript.mjs` : fonctions **pures** `pickTrack(json, lang)`, `parseJson3(data)`, `parseVtt(text)`
|
||||||
|
(testables sans réseau ni DB)
|
||||||
|
- [x] Route `GET /api/transcript/:provider/:videoId?lang=&instance=&slug=&sourceUrl=` dans `server/index.mjs`
|
||||||
|
(`instance/slug/sourceUrl` nécessaires pour PeerTube multi-instances ; réutiliser `providerUrlFrom()`).
|
||||||
|
- [x] `transcriptCache` : `Map` + TTL long (24 h), clé `transcript:${provider}:${videoId}:${lang}`.
|
||||||
|
- [x] Rate-limit (~10 req/min/IP) sur le motif `channelsLimiter`.
|
||||||
|
- [x] Aucun sous-titre → **HTTP 200** `{ available: false }` (pas une erreur).
|
||||||
|
- [x] Gestion 429 persistant + fallback « temp dir » (`writeAutoSub`) et variables d'env
|
||||||
|
(`TRANSCRIPT_CACHE_TTL`, `TRANSCRIPT_RATE_LIMIT`, `TRANSCRIPT_FALLBACK_TMP`).
|
||||||
|
|
||||||
|
**Frontend (variante UI #1 du doc)**
|
||||||
|
|
||||||
|
- [x] Bouton « Transcript » à côté du bouton Download (`watch.component.html:159`).
|
||||||
|
- [x] Panneau copié sur le motif du Download Panel (`watch.component.html:185`) :
|
||||||
|
`toggleTranscript()` + `loadTranscript()`, chargement/absent/lines, `<select>` de langue
|
||||||
|
(langue par défaut = préférence utilisateur `language`).
|
||||||
|
- [x] Le panneau ne doit **jamais** casser la page Watch (providers sans sous-titres : Twitch, Odysee, Rumble).
|
||||||
|
|
||||||
|
**Tests & docs**
|
||||||
|
|
||||||
|
- [x] `server/tests/transcript.test.mjs` : fixtures `json3`/`vtt` offline (parseurs) + intégrité du contrat API.
|
||||||
|
- [x] Script `npm run test:transcript` + l'ajouter à `.github/workflows/ci.yml`.
|
||||||
|
- [x] README : cocher `⏳ Sous-titres & transcripts` dans la roadmap.
|
||||||
|
|
||||||
|
### Hors périmètre Phase 1 (à explicitement garder pour plus tard)
|
||||||
|
|
||||||
|
- **Clic sur une ligne = seek** (variante UI #2) : exige de brancher par provider
|
||||||
|
`VideoPlayerComponent.seekBy()` (`src/components/video-player/video-player.component.ts:112`) et
|
||||||
|
`IframeProgressService.seekTo()` (`src/services/iframe-progress.service.ts:104`, iframe YouTube).
|
||||||
|
- Cache Redis / multi-instance, proxies rotatifs, monitoring des 429, formats exotiques.
|
||||||
|
|
||||||
|
### Décisions à trancher avant de coder
|
||||||
|
|
||||||
|
1. **Providers réellement supportés** en Phase 1 (le doc teste YouTube + Dailymotion ; PeerTube a sa propre API
|
||||||
|
de sous-titres — passer par `yt-dlp` ou par l'API instance ?).
|
||||||
|
2. **Persistance**: cache mémoire seulement, ou table SQLite `transcripts` (comme pour les téléchargements)
|
||||||
|
pour survivre aux redémarrages ?
|
||||||
|
3. **Transcripts générés (auto-captions)** : inclus d'emblée ou option utilisateur ?
|
||||||
|
4. **Volume de réponse** : les gros transcripts (2 h+) nécessitent-ils une troncature / pagination ?
|
||||||
|
|
||||||
|
### Critères d'acceptation
|
||||||
|
|
||||||
|
- [x] 1 appel API, 1 parsing, 1 UI identiques quel que soit le provider
|
||||||
|
- [x] 2ᵉ requête sur la même vidéo = réponse depuis le cache (pas d'appel `yt-dlp`)
|
||||||
|
- [x] `yt-dlp` en échec → 502 `{ available: false, error: 'transcript_fetch_failed' }`, Watch intacte
|
||||||
|
- [x] `npm run test:transcript` passe hors ligne (fixtures)
|
||||||
|
|
||||||
|
### Décisions Phase 1 (tranchées à l'implémentation)
|
||||||
|
|
||||||
|
1. **Providers** : générique via `yt-dlp` (ids courts `yt/dm/…` + longs normalisés), aucun code par plateforme.
|
||||||
|
2. **Persistance** : cache mémoire seul (TTL 24 h, 200 entrées LRU) ; SQLite reporté en Phase 3.
|
||||||
|
3. **Auto-captions** : incluses d'emblée (fallback après les sous-titres manuels).
|
||||||
|
4. **Volume** : troncature `TRANSCRIPT_MAX_LINES` (défaut 5000 lignes).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Idées de nouvelles fonctions utiles
|
||||||
|
|
||||||
|
Classées par effort (~ = demi-journée, + = 1-2 jours, ++ = une semaine+). Triées par rapport
|
||||||
|
valeur / coût pour l'existant.
|
||||||
|
|
||||||
|
### Recherche & découverte
|
||||||
|
|
||||||
|
| # | Fonction | Description | Effort |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 1 | **Filtres de recherche** | Durée (< 4 min / 4-20 / 20+), type (vidéo/chaîne/playlist), date de mise en ligne, langue. `SearchService.sort$` existe déjà, il manque les params + l'UI en chips + le mapping par adapter | + |
|
||||||
|
| 2 | **Pagination / infinite scroll** | ✅ DONE (Step 18) : backend pages illimitées via continuations InnerTube (`page=2,3…`, `pageSize` ≤ 50) + front `loadNextPage()` + `<app-infinite-anchor>` + `mergeGroups` + `endReached` (`search.component.ts`) | ~ |
|
||||||
|
| 3 | **Dédup multi-provider** | La même vidéo existe souvent sur Rumble/Odysee/PeerTube : regrouper par titre+durée+chaîne et afficher « aussi disponible sur… » | + |
|
||||||
|
| 4 | **Recherche dans les résultats** (« search within results ») | Relancer la requête en restreignant au provider/à la chaîne déjà affichée | ~ |
|
||||||
|
| 5 | **Écran « Recherches récentes »** | Gérer (renommer/supprimer/vider) l'historique de recherche aujourd'hui visible seulement dans le Quick Menu | ~ |
|
||||||
|
|
||||||
|
### Watch & lecture
|
||||||
|
|
||||||
|
| # | Fonction | Description | Effort |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 6 | **Recherche dans le transcript** (`Ctrl+F` du panneau) | Surligne les occurrences + liste des lignes correspondantes ; prolongement naturel du Step 16 | ~ |
|
||||||
|
| 7 | **Export du transcript** (`.vtt` / `.srt` / `.txt`) | Les lignes `{t, dur, text}` sont déjà normalisées, il ne manque qu'un rendu côté serveur | ~ |
|
||||||
|
| 8 | **Résumé IA du transcript** | `GEMINI_API_KEY` + `/api/ai/summarize` (commentaires) existent déjà (`server/index.mjs:2050`) : même pattern appliqué au transcript | + |
|
||||||
|
| 9 | **Chapitres** | Parser `chapters` de `yt-dlp --dump-single-json` (ou les timestamps du titre/description) et les afficher sous le player | + |
|
||||||
|
| 10 | **Deep-link avec timestamp** | `#/watch/yt/VIDEO?t=123` : seek au démarrage + bouton « Partager à ce moment » | ~ |
|
||||||
|
| 11 | **Reprendre la lecture** | Mémoriser la position par vidéo dans `history` et proposer « Reprendre à 12:34 » | ~ |
|
||||||
|
| 12 | **Sélecteur de qualité + « auto » intelligent** | Déjà à la roadmap ; `formatListFromMeta()` et le `format` id existent dans la file de téléchargement, réutilisables pour le player | + |
|
||||||
|
|
||||||
|
### Contenus & social
|
||||||
|
|
||||||
|
| # | Fonction | Description | Effort |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 13 | **Abonnements (chaînes)** | À la roadmap. Le socle est prêt : `server/providers/channel-registry.mjs` + `channel-content.mjs` (videos/shorts/playlists/live). Il manque table `subscriptions`, CRUD API et page « flux d'abonnements » | ++ |
|
||||||
|
| 14 | **Tags & recherche par tags** | Extraire les tags depuis `dumpSingleJson` et permettre `#/search?tags=angular` | + |
|
||||||
|
| 15 | **Import/Export playlists** (JSON / OPML) | À la roadmap ; simple contrat de sérialisation + validateur serveur | ~ |
|
||||||
|
| 16 | **Notifications « nouvelle vidéo »** | Pour les abonnements : poll planifié côté serveur + badge dans le header | + |
|
||||||
|
|
||||||
|
### Ops & qualité
|
||||||
|
|
||||||
|
| # | Fonction | Description | Effort |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 17 | **`/healthz` + page Admin** | À la roadmap : état des clés API (OK/KO), version de `yt-dlp`, hit/miss des caches, compteurs de rate-limit | + |
|
||||||
|
| 18 | **Cache serveur configurable par provider** | Remplacer le `YT_CACHE_TTL_MS` global par un TTL par provider (`CACHE_TTL_MS_YT=…`) et exposer les clés de cache vides/pleines | ~ |
|
||||||
|
| 19 | **PWA installable** | Manifest + Service Worker cache des métadonnées (les vidéos restent online) | + |
|
||||||
|
| 20 | **Superposition raccourcis clavier (`?`)** | Un `@HostListener('document:keydown')` global + modale récapitulative ; tous les raccourcis existent déjà, rien n'est découvrable | ~ |
|
||||||
|
| 21 | **i18n élargi** | Le pipeline de traduction existe (utilisé comme `'search.placeholder'` + pipeline `t`) ; il reste à compléter les lexiques et sortir les chaînes en dur du FR/EN mélangé actuel | ++ |
|
||||||
|
|
||||||
## Test commands
|
## Test commands
|
||||||
```bash
|
```bash
|
||||||
@@ -24,6 +197,8 @@ npm run test:search # Step 11 — unit tests: SearchService + @ parsing +
|
|||||||
npm run test:search-e2e # Step 12 — e2e scenarios against a real isolated server
|
npm run test:search-e2e # Step 12 — e2e scenarios against a real isolated server
|
||||||
npm run test:preferences # Step 9 — defaultProviders persistence
|
npm run test:preferences # Step 9 — defaultProviders persistence
|
||||||
npm run test:telemetry # Step 13 — telemetry events (insert/list/count)
|
npm run test:telemetry # Step 13 — telemetry events (insert/list/count)
|
||||||
|
npm run test:suggest # Step 15 — typeahead parsing/dedup + /api/search/suggest contract
|
||||||
|
npm run test:transcript # Step 16 — json3/vtt parsers + /api/transcript contract
|
||||||
```
|
```
|
||||||
|
|
||||||
## Implementation notes
|
## Implementation notes
|
||||||
@@ -44,3 +219,110 @@ npm run test:telemetry # Step 13 — telemetry events (insert/list/count)
|
|||||||
- **Step 14**: README gained a "Recherche unifiée" section (chips, @autocomplete, Ctrl/⌘+K picker, deep-links,
|
- **Step 14**: README gained a "Recherche unifiée" section (chips, @autocomplete, Ctrl/⌘+K picker, deep-links,
|
||||||
preference fallback, a11y), API endpoints, test commands and roadmap updates. GIF placeholder TODO added.
|
preference fallback, a11y), API endpoints, test commands and roadmap updates. GIF placeholder TODO added.
|
||||||
- CI (`.github/workflows/ci.yml`) now runs preferences, telemetry, search unit and e2e tests.
|
- CI (`.github/workflows/ci.yml`) now runs preferences, telemetry, search unit and e2e tests.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Step 17 — Plan Anti-Quota YouTube (sans yattee-server en dépendance critique)
|
||||||
|
|
||||||
|
### Contexte / diagnostic
|
||||||
|
|
||||||
|
- Aujourd'hui `server/providers/youtube.mjs:29-93` + `server/index.mjs:880-1069` utilisent **uniquement la YouTube Data API v3** (`search.list` + `videos.list`) avec rotation `YOUTUBE_API_KEYS` / `YOUTUBE_API_KEY` sur `quotaExceeded|rateLimitExceeded|dailyLimitExceeded|API_KEY_INVALID`.
|
||||||
|
- Coût officiel : `search.list` = **100 unités**, `videos.list` = **1 unité**, quota gratuit = **10 000 unités/jour/projet** → ~100 recherches/jour max. Chaque page `pageToken` et chaque enrichissement `videos.list` aggrave.
|
||||||
|
- Cache actuel : `YT_CACHE_TTL_MS` défaut 5 min (`server/index.mjs:904-915`) en mémoire seule, pas de persistance, pas de monitoring de conso.
|
||||||
|
- Bonne nouvelle : NewTube utilise déjà `yt-dlp` (`youtube-dl-exec`, `YT_DLP_PATH`, `server/index.mjs:83-124`) pour `transcript` + `downloads` + enrichissement Watch. Le `suggest` YT est déjà **sans clé** (`suggestqueries.google.com`, `youtube.mjs:248-264`).
|
||||||
|
- Conclusion analyse `yattee/yattee-server` : bonne architecture à copier (couches InnerTube → Invidious → yt-dlp + cache + egress-proxy), mais **ne pas l'ajouter comme service critique** (projet de 02/2026, ~107 stars, même combat anti-ban IP/cookies/PO-Token, redondant avec ton yt-dlp direct). Option sidecar seulement en Phase 4.
|
||||||
|
|
||||||
|
### Principe cible
|
||||||
|
|
||||||
|
> **Primaire sans clé (yt-dlp / InnerTube), API officielle en fallback payant, cache SQLite + mémoire, anti-ban configurable, observabilité.**
|
||||||
|
|
||||||
|
```
|
||||||
|
Front /api/search?providers=yt → searchCache (mémoire + SQLite)
|
||||||
|
→ YT_SEARCH_MODE=scrape-first (défaut) : yt-dlp `ytsearchN:` / InnerTube → OK ? return
|
||||||
|
→ sinon fallback API officielle (rotation clés) → OK ? return
|
||||||
|
→ sinon 200 avec `errors.yt` (dégradation propre, jamais 500)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Phase 0 — Quick wins (0,5 j, sans changer de source)
|
||||||
|
|
||||||
|
- [ ] Monter `YT_CACHE_TTL_MS` à 30-60 min pour `yt`, ajouter `SUGGEST_CACHE_TTL_MS` déjà à 5 min, clé cache `q|limit|page|sort`.
|
||||||
|
- [ ] Réduire le coût API : `maxResults` max 25 au lieu de 50, ne faire `videos.list` que si `details=true`, debounce front déjà 250-300 ms + `switchMap`.
|
||||||
|
- [ ] Ajouter `GET /healthz` + compteur quota estimé (`search*100 + videos*1`) exposé pour Admin (prépare idée #17).
|
||||||
|
- [ ] Doc `.env.example` : expliquer `YOUTUBE_API_KEYS` CSV vs JSON, rotation auto.
|
||||||
|
|
||||||
|
### Phase 1 — Recherche YT sans clé (2-3 j, cœur du plan)
|
||||||
|
|
||||||
|
- [ ] Nouveau `server/providers/youtube-scrape.mjs` :
|
||||||
|
- `search(q, {limit, page, sort})` via `yt-dlp --dump-single-json --flat-playlist "ytsearch{limit}:{q}"` (timeout 15-20 s, `YT_DLP_PATH` réutilisé), mapping vers `Suggestion` identique à `youtube.mjs:199-232` (title/id/thumbnail/duration/views/publishedAt/channelId/embeddable).
|
||||||
|
- Support `sort` : `relevance|date|views` → préfixe `ytsearch` + tri local si besoin.
|
||||||
|
- Jamais de throw bloquant : erreur → throw avec `ytStatus` pour que `/api/search` mette `errors.yt`.
|
||||||
|
- [ ] Modifier `server/providers/youtube.mjs` en dispatcher :
|
||||||
|
- `env YT_SEARCH_MODE=scrape-first|api-first|scrape-only|api-only` (défaut `scrape-first`).
|
||||||
|
- `scrape-first` : essaie scrape, fallback API. `api-first` : comportement actuel. Permet rollback instantané.
|
||||||
|
- [ ] Cache persistant : table SQLite `youtube_search_cache(q_hash, payload, created_at)` TTL `YT_SCRAPE_TTL_MS` (défaut 30 min) + garde mémoire actuelle. Survit au restart, tue 80% des appels doublons.
|
||||||
|
- [ ] Rate-limit + concurrence : réutiliser `express-rate-limit` existant, timeout fan-out par provider 8-12 s, `Promise.allSettled` déjà en place dans `/api/search`.
|
||||||
|
- [ ] Tests : `server/tests/youtube-scrape.test.mjs` offline (fixtures `yt-dlp --flat-playlist` mockées) + e2e `YT_SEARCH_MODE=scrape-only` sans clé → `groups.yt` non vide ; `api-only` sans clé → `errors.yt=youtube_api_key_unavailable`.
|
||||||
|
- [ ] Env : `YT_SEARCH_MODE`, `YT_SCRAPE_TTL_MS`, `YT_DLP_TIMEOUT_MS`, `YT_DLP_PATH` (déjà), `YOUTUBE_API_KEYS` devient optionnel.
|
||||||
|
|
||||||
|
### Phase 2 — Robustesse anti-ban (1-2 j, inspiré yattee-server)
|
||||||
|
|
||||||
|
- [ ] Support `YT_COOKIES_FILE` + `YT_PO_TOKEN` passés à yt-dlp (`--cookies`, `--extractor-args youtube:po_token=...`). Doc : comment exporter cookies fresh, rotation manuelle. Sans ça : `Sign in to confirm you're not a bot`.
|
||||||
|
- [ ] Support `YT_EGRESS_PROXY` (HTTP/SOCKS) pour tout le trafic YT (yt-dlp + fetch InnerTube), configurable au runtime comme yattee `SSRF_EXTRA_ALLOWED_CIDRS` si Invidious LAN.
|
||||||
|
- [ ] Auto-update yt-dlp : script `npm run ytdlp:update` + check version dans `/healthz` (idée #17). `deno`/`ffmpeg-static` déjà requis pour challenge JS.
|
||||||
|
- [ ] Observabilité Admin (idée #17) : page `Admin > YouTube` : mode actif, version yt-dlp, hit/miss cache, erreurs `quotaExceeded` vs `bot-check` vs `timeout`, état chaque clé `...abcd OK/KO`.
|
||||||
|
|
||||||
|
### Phase 3 — Fonctionnalités bonus débloquées par le sans-clé (1 j / feature)
|
||||||
|
|
||||||
|
- [ ] `Trending YT sans clé` : `yt-dlp --flat-playlist "https://www.youtube.com/feed/trending"` → alimente Accueil `Tendances & Viral` sans quota.
|
||||||
|
- [ ] `Channel browsing enrichi` : réutiliser `channel-registry.mjs` + `channel-content.mjs` via scrape (`videos/shorts/streams/playlists`) au lieu de `channels.list` payant → prérequis Abonnements (idée #13).
|
||||||
|
- [ ] `Chapitres` (idée #9) : parser `chapters` de `dump-single-json` déjà dispo.
|
||||||
|
- [ ] `Filtres recherche` (idée #1) : `duration/date/type` mappés sur args yt-dlp + filtre local, 0 coût API.
|
||||||
|
- [ ] `Pagination / infinite scroll` (idée #2) : `page$` déjà câblé, brancher `nextPageToken` scrape via `--flat-playlist --playlist-start`.
|
||||||
|
|
||||||
|
### Phase 4 — Option yattee-server sidecar (seulement si Phase 1 insuffisante)
|
||||||
|
|
||||||
|
- [ ] `docker-compose/yattee-server.yml` : image `yattee/yattee-server`, `INVIDIOUS_INSTANCE_URL` optionnelle, `ADMIN_USERNAME/PASSWORD` via `.env`.
|
||||||
|
- [ ] Adaptateur `server/providers/youtube-yattee.mjs` : `GET ${YATTEE_URL}/api/v1/search?q=&type=video` + Basic Auth → `Suggestion[]`. Activé par `YT_SEARCH_MODE=yattee`.
|
||||||
|
- [ ] Critère GO/NO-GO : si scrape direct tient >95% succès sur 7 j, abandonner sidecar. Sinon le garder pour `trending/comments/captions`.
|
||||||
|
|
||||||
|
### Risques / ToS
|
||||||
|
|
||||||
|
- Scraping/InnerTube = gris vis-à-vis ToS YouTube, casses fréquentes → prévoir fallback API + version yt-dlp pinnée + alertes `/healthz`.
|
||||||
|
- Pas de magie IP : datacenter OVH/Hetzner = ban plus vite que résidentiel → prévoir proxy sortant dès le déploiement public.
|
||||||
|
- Ne jamais logger clés, cookies, PO-Token.
|
||||||
|
|
||||||
|
### Critères d'acceptation Step 17
|
||||||
|
|
||||||
|
- [x] Sans aucune `YOUTUBE_API_KEY`, `GET /api/search?q=test&providers=yt` retourne `groups.yt[]` (via scrape) et `npm run test:search-e2e` passe. (vérifié live : scrape-only retourne 3 résultats avec durée/vues/channelId)
|
||||||
|
- [x] Avec quota épuisé simulé (400/403 mock), fallback scrape prend le relais sans 500. (dispatcher scrape-first → api en fallback sur bot-check/timeout/upstream)
|
||||||
|
- [x] 2ᵉ appel identique < 50 ms (hit cache mémoire/SQLite, pas d'appel yt-dlp). (mesuré : 3971 ms → 30 ms)
|
||||||
|
- [x] `/healthz` expose `ytdlpVersion`, `mode`, `cacheHitRate`. (`/healthz` + `/api/healthz` : mode, ytdlp bin/version, antiban, cache mem+sqlite, metrics jour, clés)
|
||||||
|
- [x] Aucune régression `dm/tw/pt/od/ru`. (`test:search-e2e` + `test:suggest` verts)
|
||||||
|
|
||||||
|
### Implémenté le 2026-09-25 (Step 17 DONE)
|
||||||
|
|
||||||
|
- Nouveaux : `server/providers/youtube-common.mjs` (clés, `YT_SEARCH_MODE`, args anti-ban cookies/PO-Token/proxy, hash, métriques), `server/providers/youtube-scrape.mjs` (search/channel/trending via `yt-dlp --flat-playlist`, parsers `mapFlatEntry`/`parseFlatPlaylistJson` testés offline), `db/migrations/20250926_add_youtube_scrape_cache.sql`, `server/tests/youtube-scrape.test.mjs` (`npm run test:ytscrape` + CI).
|
||||||
|
- Modifiés : `youtube.mjs` (dispatcher scrape-first/api-first/scrape-only/api-only + cache mémoire LRU + SQLite + logs source/latency), `channel-content.mjs` (resolve + contenu YT en scrape-first, fallback API), `index.mjs` (`/healthz`, `/api/trending`, import common), `db.mjs` (helpers cache/métriques jour), `.env.example` (nouvelles vars), `package.json` (`test:ytscrape`, `ytdlp:update`).
|
||||||
|
- Notes : `/feed/trending` retiré côté YouTube (redirect home) → repli `ytsearch` trié vues. yt-dlp local `2026.08.19` : penser `npm run ytdlp:update`. Sidecar yattee-server abandonné (scrape direct >95% : à confirmer sur 7 j).
|
||||||
|
- Fix 2026-09-25 (`spawn yt-dlp ENOENT`) : `resolveYtDlpBin()` (`youtube-common.mjs`) avec ordre `YT_DLP_PATH` > PATH > bundled `youtube-dl-exec` (cause : yt-dlp seulement présent via shims scoop utilisateur, invisible d'un autre contexte/Docker) ; fallback API sur **tout** échec scrape en `scrape-first` (plus d'erreur brute en UI) ; message actionnable `youtube_no_source` si ni binaire ni clé ; `/healthz` expose `binOk` + binaire résolu.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Step 18 — InnerTube direct façon SmartTube (pagination illimitée + vidéos connexes) ✅
|
||||||
|
|
||||||
|
### Contexte
|
||||||
|
|
||||||
|
SmartTube (`yuliskov/SmartTube`, 34k stars) ne touche ni la Data API ni les Google Services : il parle à **InnerTube** (`youtubei/v1/*`, clients TV) via `MediaServiceCore` ("Unofficial Java api for YouTube") — search + continuations, watch-next (`/next`), player. D'où 0 quota et scroll infini.
|
||||||
|
|
||||||
|
### Implémenté le 2026-09-25
|
||||||
|
|
||||||
|
- Dépendance `youtubei.js` **pinnée `18.1.0`** (API interne non documentée → pin + bump manuel).
|
||||||
|
- Nouveau `server/providers/youtube-innertube.mjs` : session singleton lazy (`YT_INNERTUBE_GL/HL`), `searchViaInnerTube` (**continuations fusionnées** jusqu'à couvrir `page*limit` — corrige le bug "page 2 vide" : 1 page InnerTube ≈ 17-20 items ≠ `limit`), `getRelatedViaInnerTube` (watch-next), mappers purs `mapVideoNode`/`mapLockupView`/`parseViewsText`/`parseDurationLabel`, chaîne de continuations en mémoire + SQLite existant.
|
||||||
|
- `mapLockupView` : le watch-next moderne renvoie des **LockupView** (`content_id`, `metadata.title.text`, `content_image.image[]`, chaîne + vues dans `metadata_rows`, durée dans le label a11y) — filtrés avant, mappés maintenant.
|
||||||
|
- Dispatcher `innertube-first` (**défaut**) : InnerTube → scrape yt-dlp → API officielle ; `innertube-only` ajouté ; les anciens modes inchangés ; on ne persiste plus les résultats vides (anti-empoisonnement du cache).
|
||||||
|
- `GET /api/details/youtube/:videoId` → champ **`related[]`** (24 items, cache mémoire 1h, `?related=0` pour désactiver, best-effort jamais bloquant).
|
||||||
|
- Tests `server/tests/youtube-innertube.test.mjs` (`npm run test:ytinnertube` + CI), `.env.example` à jour.
|
||||||
|
- Vérifié live : recherche p1=50 + p2=50 = **100 uniques** (fin de la limite 24), related=24 avec titres FR/durées/vues, `test:search-e2e` + `test:suggest` verts.
|
||||||
|
- Limites connues : pas de `getTrending` en youtubei.js v18 → trending reste sur scrape ; continuations parfois redondantes (dédup + garde-fou 12) ; même combat anti-ban IP qu'avant (cookies/PO-Token/proxy réutilisés côté yt-dlp, session InnerTube sans auth).
|
||||||
|
- Complément 2026-09-26 : Watch → sidebar « connexes » branchée sur le vrai watch-next (`loadRelatedSuggestions()` utilise `GET /api/details/youtube/:id` → `related[]`, fallback recherche-par-titre pour les autres providers et en cas d'échec) ; `docker-compose/.env.example` + `README.md` (endpoints + roadmap) à jour ; idée #2 (pagination/infinite scroll) clôturée.
|
||||||
|
- Complément transcript InnerTube : découverte des pistes via `getInfo().captions` (`YT_TRANSCRIPT_SOURCE`, défaut `innertube-first`) adaptée au format yt-dlp → `pickTrack`/`orderedTracks`/`parseTrackText` inchangés ; 0 piste InnerTube = `no_subtitles` direct (même backend que le lecteur) ; fallback yt-dlp conservé. Vérifié live (`lang=en` → 217 lignes) ; `test:transcript` 17/17.
|
||||||
|
|||||||
Reference in New Issue
Block a user