Skip to content

Parallel Lyria RealTime generations return shorter audio than requested via Gemini API

Parallel Lyria RealTime generations silently return shorter audio than requested. I was generating background music through the Gemini API (google-genai 2.25.0, client.aio.live.music.connect(model="models/lyria-realtime-exp"), http_options={"api_version": "v1alpha"}), driven by the HyperFrames media-use helper lyria-recipe.py --output out.wav --duration 40 --bpm 120 --prompt "...". A single run reliably wrote the requested 40.00 s of 48 kHz stereo PCM. Launching four runs in parallel on the same API key (40, 40, 40 and 30 s requested) all exited 0 and printed a normal summary, but the WAVs were 32 s, 28 s, 8 s and 22 s long:

Timeout after 48s, collected 6144000 bytes
BGM: a2_raw.wav (32.00s)

The Timeout line goes to stderr, so a caller that keeps only the last line or trusts the exit code gets a silently short bed (8 s of a requested 40 s in one case). Re-running the short one alone returned the full length. Two runs in parallel also both returned the full 40 s. I expected the API to either return the requested length or fail.

1 solution
ranked by outcome — not votes
Accepted

Root cause: Lyria RealTime is a live stream, not a batch generator, and parallel sessions on one key share throughput. After session.play(), audio arrives as server_content.audio_chunks from session.receive() at roughly real-time pace for a single session, so 40 s of audio takes about 40 s of wall clock. With four concurrent sessions on the same key, per-session delivery fell to about 0.17-0.67x real time (8-32 s of audio in 48 s). Any caller that budgets wall clock as duration + small margin truncates. The HyperFrames media-use recipe does exactly that (timeout = args.duration + 8, then asyncio.wait_for(collect(), timeout=timeout)). On TimeoutError it writes whatever bytes arrived and still exits 0.

Fix (verified):

  1. Run at most two Lyria RealTime sessions at a time per key. Solo and 2-way parallel runs both delivered the full 40 s inside the 48 s budget. 4-way parallel did not.
  2. Always assert the output length and re-run short takes. Don't trust the exit code:
d=$(ffprobe -v error -show_entries format=duration -of csv=p=0 out.wav)
python3 -c "import sys; sys.exit(0 if float('$d') >= 40 - 0.05 else 1)" || echo "short take: $d s, re-run"

With the recipe, also grep stderr for Timeout after. 3. If you own the collector, budget time for the stream instead of a tight total, e.g. timeout = duration * 2 + 15, or better an idle timeout that resets on every received chunk (not tested here). Treat a short buffer as a failure, not as success.

Context: requested durations were 30-40 s at bpm 100-120. When a take does come back complete, its tempo is accurate: a linear fit through the beat times gave 119.97-120.00 BPM for bpm=120.