Skip to content

Performance & hardware

ConversationSimulator runs a local LLM, optionally local STT (Whisper) and TTS (Kokoro), all on the same machine as the app. Performance depends heavily on available hardware. The tiers below are product targets, not guarantees — actual results vary by model family, driver versions, and system load.

These are the official latency budgets for the mid-spec reference machine (Apple M2 / NVIDIA RTX 3060 equivalent). In-app performance warnings fire when a measurement exceeds the corresponding budget. Nightly CI smoke tests flag any regression greater than 20 % against these values.

MetricBudgetCondition
Cold start → interactive Home< 10 sModel already downloaded; startup to playable state
Time-to-first-token (TTFT)< 2.5 sStarter model (4 B Q4_K_M) on recommended-tier hardware
TTS first-audio chunk< 1.5 sKokoro sidecar, sentence-level streaming
STT round-trip< 2 s10-word utterance, whisper.cpp small model

All values in the LATENCY_BUDGETS constant in packages/shared/src/types/metrics.ts. PerformanceWarning thresholds are derived from these constants.

TierExample hardwareExpected NPC first-token latencyNotes
FastApple M-series (≥M2), NVIDIA RTX 3060+< 1 sGPU offload enabled; small–medium models
ComfortableOlder integrated GPU, 6-core CPU1–2.5 sReduced GPU layers or CPU-only; small model recommended
SlowLow-power CPU, 4 GB RAM2.5–10 sText-only recommended; disable VAD and TTS
Unsupported< 8 GB RAM total, no GPU> 10 s or OOMCannot run local models reliably

Note: “Slow” tier is functional but will trigger in-app performance warnings that link to Runtime Settings.

Measured with the starter model (Qwen3 4B Q4_K_M, 2.5 GB, context 8192) using a 30-token scripted turn. All timings in seconds (median of 5 runs). STT and TTS use whisper.cpp small and Kokoro TTS respectively.

MetricFast tier (M2 / RTX 3060)Comfortable tier (6-core CPU + iGPU)Slow tier (4-core CPU, no GPU)
Cold start → Home~3 s~6 s~9 s
Time-to-first-token~0.8 s~1.8 s~2.4 s
Full response (30 tok)~1.5 s~3.5 s~8 s
TTS first audio~0.5 s~1.0 s~1.4 s
STT round-trip~0.6 s~1.2 s~1.9 s

Budget column for reference: TTFT < 2.5 s, TTS < 1.5 s, STT < 2 s, cold start < 10 s. Comfortable and Slow tiers operate within budget on the starter model; high-quality models (14 B+) push into Slow or Unsupported territory on those tiers.

These results are for the Steam system-requirements reference platform described in #283’s docs. For players whose hardware matches a tier, the table sets honest download and runtime expectations.

The app tracks the following timings locally. No data leaves your machine.

MetricDescriptionWarning threshold
Session startTime from start request to NPC opening> 10 s
First tokenTime from player turn to first streamed NPC token> 2.5 s
Full responseTime from player turn to complete NPC response> 10 s
STT finalTime for speech-to-text to return a transcript> 2 s
TTS first sentenceTime for first audio chunk to be ready> 1.5 s
Debrief generationTime for debrief to generate—

Conversation-screen metrics (session start, first token, full response, STT final, TTS first sentence) are visible in the Developer debug panel; debrief-generation latency is shown on the debrief screen. Both require dev mode (enable it in Settings).

The conversation screen also remembers how long recent turns took, so it can tell the player what to expect from the next one — see How long a turn will take.

A local model offers the app no progress signal while it works — no percentage, no token count to divide — so the conversation screen answers “how long will this take?” with the only honest evidence it has: how long this machine’s own recent turns took.

While the NPC is thinking, the screen shows a clock and a bar:

What you seeWhat it means
22s over a hatched barNothing timed yet on this model. The clock still runs, so a long wait never looks like a frozen app.
22s / ~40s over a filling barA turn is expected to cost about 40 s on this machine, and the caption says roughly how much of it is left.
1m 20s / ~40s in amberThe turn has outrun its estimate. The bar stops pretending to know, and says the reply is not lost.

The panel is the screen’s status for as long as the NPC is out, so nothing else repeats it: its heading switches from “NPC is thinking” to “NPC is replying” once the reply starts arriving word by word, and the whole panel disappears the moment the reply is on screen. A turn is not quite over at that point — the app is still collecting the state changes that went with the reply, which it says plainly (“Finishing the turn…”) while the input stays disabled — and that tail is included in the duration it files away, so the estimate it quotes next time is for the whole round trip.

Worth knowing about the number it quotes:

  • It is the median of up to eight recent turns, not the mean. One turn that stalled because the system paged the model back in should not inflate the next five.
  • Timings are kept per model. Switching to a smaller model — the app’s own advice when turns are slow — starts the estimate over instead of quoting the old model’s minutes at the new one.
  • They never leave the machine. The samples are bare durations in milliseconds, held in browser local storage under convsim.turnTiming, for one model at a time. Clear all local data in Settings forgets them along with everything else.
  • A turn recovered by polling is not timed. Past the five-minute deadline the app finds the reply by asking the session what it recorded (see Timeout errors), which measures when the screen noticed the turn rather than what the model spent on it.
  • The bar never fills completely while a reply is still out. A full bar with nothing on screen reads as a turn the app lost.

Estimates appear from the second turn of a fresh install: the first one has nothing to go on, so it only shows the clock.

The app surfaces actionable warnings when thresholds are exceeded. Each links to Runtime Settings.

NPC response is slow (first token > 2.5 s)

Section titled “NPC response is slow (first token > 2.5 s)”
  • Try a smaller model. A 4 B-parameter model runs 2–4× faster than an 8 B model on the same hardware. See the Model Manager for size-categorised options.
  • Increase GPU layers. If you have a GPU, set GPU Layers to -1 (all layers to GPU) in Runtime Settings.
  • Reduce context length. Shorter context = faster KV-cache fill. Try 2048 or 4096 instead of the model default.

All of the above, plus:

  • Switch to push-to-talk. VAD (hands-free) adds latency before each turn. Push-to-talk removes that overhead.
  • Switch to text-only mode. Skip STT/TTS entirely for the lowest-latency experience.
  • Disable TTS. Uncheck “Enable TTS” in scenario setup. Text is delivered immediately.
  • The Kokoro TTS sidecar requires a live connection to the sidecar process. If it is slow, check that the process is running and has sufficient CPU resources.

STT recognition is slow (round-trip > 2 s)

Section titled “STT recognition is slow (round-trip > 2 s)”
  • Switch to push-to-talk. VAD silence detection adds overhead. Push-to-talk avoids it.
  • Switch to text input. Eliminates STT entirely; type your player turns instead.

When whisper.cpp is not running or returns an error, the app falls back to text input automatically. A status indicator appears at the top of the conversation screen.

When Kokoro is not running, the app operates in text-only mode. No audio is played. The scenario plays normally; only the voice output is absent.

A slow turn is not a failed turn. A local model on CPU-only hardware can spend a minute or more on a single reply — mostly reading the prompt back in — and the app waits for it. The clock and bar described in How long a turn will take are on screen for the whole wait; after five seconds the app adds that the NPC is taking longer than usual, and after thirty seconds that the reply has not been thrown away.

Two separate limits can still end a turn early:

LimitValueWhat happens
Engine went quiet180 s with nothing sent (CONVSIM_LLAMA_CPP_CHAT_TIMEOUT) — this covers reading the prompt back in as well as the gaps between wordsThe turn fails with a timeout error. Nothing is recorded, so you can retype the same turn.
App gave up waiting10 minFrom 5 minutes on, the app stops trusting the request and starts asking the session what it actually recorded, every 15 seconds. As soon as the reply has landed it is shown and play continues. If the request itself answers while that is going on — with the reply, or with a failure of its own — that answer is used straight away. Only if nothing has landed by 10 minutes does the turn fail with a timeout error. (A session started with transcript saving off has nothing to ask about, so the app simply keeps waiting on the request until then.)

The session is not ended by either case — you can retry the same turn. The error message includes the same suggestions listed above.

Turning off features to improve performance

Section titled “Turning off features to improve performance”

All of the following can be changed without restarting a session:

FeatureHow to disablePerformance impact
TTSUncheck “Enable TTS” in scenario setupRemoves Kokoro synthesis latency
VADSwitch to push-to-talk in scenario setupRemoves silence-detection overhead
State metersUncheck “Show state meters” in setupMinor
Transcript savingUncheck “Save transcript” in setupMinor

Runtime Settings (GPU layers, context length, threads, temperature) take effect on the next model load.