docs: audio workspace tab design (2-phase: extract/edit/classify + dataset)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,155 @@
|
||||
# Audio workspace tab — Design
|
||||
|
||||
**Goal:** Promote audio from a pair of cramped transport-row controls to a
|
||||
first-class **Audio** deck tab that unifies everything audio: exact-length
|
||||
extraction, light non-destructive editing of the extracted sound, an
|
||||
audio-dedicated dataset workflow, and the existing scan/classify controls.
|
||||
|
||||
**Runs in:** Python/Qt client (`main.py`) + `core/` helpers. No server changes.
|
||||
|
||||
**Approach:** Replace the current **Scan** deck tab with an **Audio** tab.
|
||||
Because the tab now covers four concerns, its body is a `QToolBox` accordion
|
||||
(one pane open at a time) rather than flat stacked group-boxes. Existing
|
||||
scan/extraction widgets are re-parented (same handlers/signals) — no behavior
|
||||
change beyond relocation.
|
||||
|
||||
**Deck order after:** `Export | Crop & Track | Audio`. The `ScanResultsPanel`
|
||||
(3rd splitter pane) is untouched.
|
||||
|
||||
Delivered in **two phases**. Phase 1 is self-contained and shippable; Phase 2
|
||||
adds the dataset workflow.
|
||||
|
||||
---
|
||||
|
||||
## Deck wiring (both phases)
|
||||
|
||||
`_build_control_deck` (main.py:4661): swap `_tab_scan` → `_tab_audio` —
|
||||
`_deck_key="audio"`, `_label="Audio"`, `_pinned=False`; replace it in
|
||||
`_deck_panels`; one `addTab`. The pin/split machinery is `_deck_panels`-driven,
|
||||
so it picks the new panel up automatically. Old persisted `"scan"` pin-state is
|
||||
ignored (harmless). Nothing auto-switches to the old scan tab — verified:
|
||||
`_tab_scan` is referenced only in construction + its builder.
|
||||
|
||||
`_build_scan_tab` is renamed `_build_audio_tab` and rebuilt as a `QToolBox`
|
||||
with the panes below.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Extract, Edit, Classify
|
||||
|
||||
### Pane 1 — Extract & Edit
|
||||
|
||||
**Relocated from the transport row** (removed there): `_spn_audio_len` (length)
|
||||
and `_btn_extract_audio` (`♪ Extract`). The teal timeline region band is
|
||||
unaffected — still driven by `_spn_audio_len`.
|
||||
|
||||
**New controls:**
|
||||
- `_cmb_audio_fmt` — output format (WAV / MP3 / FLAC / M4A / OGG / OPUS),
|
||||
persisted to QSettings `audio_extract_fmt`; pre-selects the save-dialog filter
|
||||
so the format isn't re-picked each time.
|
||||
- **Waveform strip** — a compact widget rendering the current region (or
|
||||
last-extracted clip). Generated from decoded PCM (reuse
|
||||
`core/audio_scan._load_audio_ffmpeg`) or `ffmpeg showwavespic`. Read-only in
|
||||
v1 (no drag handles).
|
||||
- **Play / Stop** — audition the current region/clip. Playback backend decided
|
||||
in the plan (likely a short ffmpeg-decoded buffer through a
|
||||
`QMediaPlayer`/`QSoundEffect`, kept separate from the mpv video player).
|
||||
|
||||
**Editing (non-destructive, ffmpeg filter chain at render time):**
|
||||
|
||||
| Control | Filter |
|
||||
|---|---|
|
||||
| Trim in / out (refine past the length spinbox) | `atrim` / adjusted `-ss` / `-t` |
|
||||
| Fade in / out (s) | `afade=t=in` / `afade=t=out` |
|
||||
| Normalize (toggle) | `loudnorm` |
|
||||
| Gain (dB) | `volume` |
|
||||
|
||||
`core/ffmpeg.build_audio_clip_command` gains an optional `filters` param →
|
||||
appended as `-af "<chain>"`. Foley/existing callers pass none → byte-identical
|
||||
output. Extract applies the current edit settings; nothing is written until you
|
||||
save.
|
||||
|
||||
### Pane 2 — Scan / Classify
|
||||
|
||||
Relocated verbatim from the old Scan tab: model combo + `_btn_model_history`;
|
||||
`_btn_scan`, `_btn_auto_export`, `_btn_speech`, `_btn_scan_mode`;
|
||||
`_spn_auto_fuse`, `_sld_threshold`. Same handlers, same `ScanResultsPanel`
|
||||
output.
|
||||
|
||||
*(Train / Hard-negatives / Dataset-stats buttons join Pane 3 in Phase 2; in
|
||||
Phase 1 they stay wherever they are today — the menu bar.)*
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — Audio dataset
|
||||
|
||||
### Pane 3 — Dataset (audio-dedicated)
|
||||
|
||||
A dedicated audio dataset **workflow + view**, combining browser-plus-label,
|
||||
add-from-extraction, and in-tab access to the existing tools.
|
||||
|
||||
**Key data-model finding (the one open decision — confirm before building):**
|
||||
the existing classifier trains on **(source_video, time-ranges)** per profile —
|
||||
`db.get_training_data(...)` returns `(source_video_path, positive_times,
|
||||
soft_times, negative_times)` per video; standalone clip *files* are not the
|
||||
training unit. An extracted audio region always knows its
|
||||
`(source_video, start, end)`. So the cheapest, most powerful integration is
|
||||
**merged**: "Add to dataset" records the region as a labeled positive/negative
|
||||
time-range in the **same** per-profile DB the video-export workflow already
|
||||
feeds — no new trainer, no schema change. The Audio dataset pane is then an
|
||||
**audio-oriented view/filter** over that shared dataset.
|
||||
|
||||
- **Recommended (merged):** audio extraction becomes another way to label
|
||||
training data; reuses `get_training_data`, `get_training_stats`,
|
||||
hard-negatives, and `train_classifier` unchanged.
|
||||
- **Alternative (separate):** a standalone folder of `.wav` clips + own
|
||||
`dataset.json` + a new file-based training path. More code, duplicates the
|
||||
trainer. Only needed if the audio set must stay entirely independent of the
|
||||
video profiles.
|
||||
|
||||
*(Brainstorming default was "separate"; inspecting the DB flips the
|
||||
recommendation to "merged." Flagged for confirmation.)*
|
||||
|
||||
**Pane contents (assuming merged):**
|
||||
- **Add to dataset** — from the current/last extraction:
|
||||
`[class ▾]` + `(•positive / ○negative)` + `[+ Add]`. Writes a labeled
|
||||
time-range into the DB for the active profile (and, for standalone-clip
|
||||
parity, writes the clip's `dataset.json` label via `upsert_clip_annotation`
|
||||
when a file was saved).
|
||||
- **Browser** — a list of the audio dataset entries (class, pos/neg, source,
|
||||
time-range), filterable by class; select to seek the source video to that
|
||||
region; remove entries.
|
||||
- **Tools** — in-tab buttons for the existing dialogs: *Dataset Statistics*
|
||||
(`DatasetStatsDialog`), *Manage Hard Negatives* (`HardNegativesDialog`),
|
||||
*Train classifier* (existing Train dialog).
|
||||
|
||||
---
|
||||
|
||||
## Persistence & migration
|
||||
|
||||
- QSettings gains `audio_extract_fmt` (+ edit defaults: fade/normalize/gain)
|
||||
alongside existing `audio_extract_len`, `audio_extract_dir`.
|
||||
- Deck pin-state key `"scan"` → `"audio"`; absent → default unpinned.
|
||||
- Phase 2 (merged): no schema change — reuses existing training tables /
|
||||
`dataset.json`. Phase 2 (separate) would add a folder convention only.
|
||||
|
||||
## Testing
|
||||
|
||||
- `tests/test_ui_structure.py`: update `test_extract_audio_controls_exist` to
|
||||
locate controls in the Audio tab; add a test that the Audio deck tab exists as
|
||||
a `QToolBox` with Extract/Classify panes and that scan controls remain
|
||||
reachable.
|
||||
- `tests/test_utils.py`: extend the audio-command tests for the new `-af`
|
||||
filter chain (trim/fade/normalize/gain), keeping existing codec-by-extension
|
||||
assertions.
|
||||
- Phase 2: a DB test that "add to dataset" produces a training-data row
|
||||
consumable by `get_training_data`.
|
||||
|
||||
## What this does NOT do
|
||||
|
||||
- No full waveform-drag editor in v1 — editing is parametric
|
||||
(trim/fade/normalize/gain).
|
||||
- No change to the Foley/LTX-2 export pipeline or `ScanResultsPanel`.
|
||||
- No new ML model or embedding change.
|
||||
- Phase 1 makes no DB/schema change; the dataset-model decision is confined to
|
||||
Phase 2.
|
||||
Reference in New Issue
Block a user