a315093743
Adds sync_strength (0.0–3.0, default 1.0) to PrismAudioSampler. The scale is applied post-conditioner (after Sync_MLP) to the conditioning tensor before it enters the DiT. Since CFG always uses zeros as the null sync embedding, this cleanly scales the sync guidance signal: effective_sync_guidance = cfg_scale * (sync_strength * cond - 0) Higher values tighten temporal audio-video alignment; 0.0 disables sync guidance entirely (audio conditioned only by video + text features). Not applied in T2A mode where sync is replaced by the learned empty_sync_feat. Also logs sync temporal coverage vs audio target duration, with a warning when they differ by more than 0.5s (stale or mismatched features). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>