- Detect CUDA version at venv creation time and install matching jax[cuda12/13]
instead of hardcoded jax[cuda13] — was broken on CUDA 12.x (most systems)
- Include fps in cache hash: same video+caption at different fps previously
returned stale cached features with wrong frame sampling
- Guard frame index lists with max(1,...)/max(8,...) to prevent torch.stack([])
crash on very short input clips; sync minimum is 8 to match Synchformer's
segment size requirement
- Remove mediapy from managed venv packages — not imported anywhere
- Warn when caption_cot is empty (produces degenerate text features)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>