Performance & hardware
ConversationSimulator runs a local LLM, optionally local STT (Whisper) and TTS (Kokoro), all on the same machine as the app. Performance depends heavily on available hardware. The tiers below are product targets, not guarantees — actual results vary by model family, driver versions, and system load.
Latency budgets
Section titled “Latency budgets”These are the official latency budgets for the mid-spec reference machine (Apple M2 / NVIDIA RTX 3060 equivalent). In-app performance warnings fire when a measurement exceeds the corresponding budget. Nightly CI smoke tests flag any regression greater than 20 % against these values.
| Metric | Budget | Condition |
|---|---|---|
| Cold start → interactive Home | < 10 s | Model already downloaded; startup to playable state |
| Time-to-first-token (TTFT) | < 2.5 s | Starter model (4 B Q4_K_M) on recommended-tier hardware |
| TTS first-audio chunk | < 1.5 s | Kokoro sidecar, sentence-level streaming |
| STT round-trip | < 2 s | 10-word utterance, whisper.cpp small model |
All values in the LATENCY_BUDGETS constant in packages/shared/src/types/metrics.ts. PerformanceWarning thresholds are derived from these constants.
Hardware tiers
Section titled “Hardware tiers”| Tier | Example hardware | Expected NPC first-token latency | Notes |
|---|---|---|---|
| Fast | Apple M-series (≥M2), NVIDIA RTX 3060+ | < 1 s | GPU offload enabled; small–medium models |
| Comfortable | Older integrated GPU, 6-core CPU | 1–2.5 s | Reduced GPU layers or CPU-only; small model recommended |
| Slow | Low-power CPU, 4 GB RAM | 2.5–10 s | Text-only recommended; disable VAD and TTS |
| Unsupported | < 8 GB RAM total, no GPU | > 10 s or OOM | Cannot run local models reliably |
Note: “Slow” tier is functional but will trigger in-app performance warnings that link to Runtime Settings.
Results table by hardware tier
Section titled “Results table by hardware tier”Measured with the starter model (Qwen3 4B Q4_K_M, 2.5 GB, context 8192) using a 30-token scripted turn. All timings in seconds (median of 5 runs). STT and TTS use whisper.cpp small and Kokoro TTS respectively.
| Metric | Fast tier (M2 / RTX 3060) | Comfortable tier (6-core CPU + iGPU) | Slow tier (4-core CPU, no GPU) |
|---|---|---|---|
| Cold start → Home | ~3 s | ~6 s | ~9 s |
| Time-to-first-token | ~0.8 s | ~1.8 s | ~2.4 s |
| Full response (30 tok) | ~1.5 s | ~3.5 s | ~8 s |
| TTS first audio | ~0.5 s | ~1.0 s | ~1.4 s |
| STT round-trip | ~0.6 s | ~1.2 s | ~1.9 s |
Budget column for reference: TTFT < 2.5 s, TTS < 1.5 s, STT < 2 s, cold start < 10 s. Comfortable and Slow tiers operate within budget on the starter model; high-quality models (14 B+) push into Slow or Unsupported territory on those tiers.
These results are for the Steam system-requirements reference platform described in #283’s docs. For players whose hardware matches a tier, the table sets honest download and runtime expectations.
What the app measures
Section titled “What the app measures”The app tracks the following timings locally. No data leaves your machine.
| Metric | Description | Warning threshold |
|---|---|---|
| Session start | Time from start request to NPC opening | > 10 s |
| First token | Time from player turn to first streamed NPC token | > 2.5 s |
| Full response | Time from player turn to complete NPC response | > 10 s |
| STT final | Time for speech-to-text to return a transcript | > 2 s |
| TTS first sentence | Time for first audio chunk to be ready | > 1.5 s |
| Debrief generation | Time for debrief to generate | — |
Conversation-screen metrics (session start, first token, full response, STT final, TTS first sentence) are visible in the Developer debug panel; debrief-generation latency is shown on the debrief screen. Both require dev mode (enable it in Settings).
The conversation screen also remembers how long recent turns took, so it can tell the player what to expect from the next one — see How long a turn will take.
How long a turn will take
Section titled “How long a turn will take”A local model offers the app no progress signal while it works — no percentage, no token count to divide — so the conversation screen answers “how long will this take?” with the only honest evidence it has: how long this machine’s own recent turns took.
While the NPC is thinking, the screen shows a clock and a bar:
| What you see | What it means |
|---|---|
22s over a hatched bar | Nothing timed yet on this model. The clock still runs, so a long wait never looks like a frozen app. |
22s / ~40s over a filling bar | A turn is expected to cost about 40 s on this machine, and the caption says roughly how much of it is left. |
1m 20s / ~40s in amber | The turn has outrun its estimate. The bar stops pretending to know, and says the reply is not lost. |
The panel is the screen’s status for as long as the NPC is out, so nothing else repeats it: its heading switches from “NPC is thinking” to “NPC is replying” once the reply starts arriving word by word, and the whole panel disappears the moment the reply is on screen. A turn is not quite over at that point — the app is still collecting the state changes that went with the reply, which it says plainly (“Finishing the turn…”) while the input stays disabled — and that tail is included in the duration it files away, so the estimate it quotes next time is for the whole round trip.
Worth knowing about the number it quotes:
- It is the median of up to eight recent turns, not the mean. One turn that stalled because the system paged the model back in should not inflate the next five.
- Timings are kept per model. Switching to a smaller model — the app’s own advice when turns are slow — starts the estimate over instead of quoting the old model’s minutes at the new one.
- They never leave the machine. The samples are bare durations in milliseconds, held in browser local storage under
convsim.turnTiming, for one model at a time. Clear all local data in Settings forgets them along with everything else. - A turn recovered by polling is not timed. Past the five-minute deadline the app finds the reply by asking the session what it recorded (see Timeout errors), which measures when the screen noticed the turn rather than what the model spent on it.
- The bar never fills completely while a reply is still out. A full bar with nothing on screen reads as a turn the app lost.
Estimates appear from the second turn of a fresh install: the first one has nothing to go on, so it only shows the clock.
What to do when the app is slow
Section titled “What to do when the app is slow”The app surfaces actionable warnings when thresholds are exceeded. Each links to Runtime Settings.
NPC response is slow (first token > 2.5 s)
Section titled “NPC response is slow (first token > 2.5 s)”- Try a smaller model. A 4 B-parameter model runs 2–4× faster than an 8 B model on the same hardware. See the Model Manager for size-categorised options.
- Increase GPU layers. If you have a GPU, set GPU Layers to
-1(all layers to GPU) in Runtime Settings. - Reduce context length. Shorter context = faster KV-cache fill. Try 2048 or 4096 instead of the model default.
Full response is very slow (> 10 s)
Section titled “Full response is very slow (> 10 s)”All of the above, plus:
- Switch to push-to-talk. VAD (hands-free) adds latency before each turn. Push-to-talk removes that overhead.
- Switch to text-only mode. Skip STT/TTS entirely for the lowest-latency experience.
TTS audio is slow (first audio > 1.5 s)
Section titled “TTS audio is slow (first audio > 1.5 s)”- Disable TTS. Uncheck “Enable TTS” in scenario setup. Text is delivered immediately.
- The Kokoro TTS sidecar requires a live connection to the sidecar process. If it is slow, check that the process is running and has sufficient CPU resources.
STT recognition is slow (round-trip > 2 s)
Section titled “STT recognition is slow (round-trip > 2 s)”- Switch to push-to-talk. VAD silence detection adds overhead. Push-to-talk avoids it.
- Switch to text input. Eliminates STT entirely; type your player turns instead.
STT unavailable
Section titled “STT unavailable”When whisper.cpp is not running or returns an error, the app falls back to text input automatically. A status indicator appears at the top of the conversation screen.
TTS unavailable
Section titled “TTS unavailable”When Kokoro is not running, the app operates in text-only mode. No audio is played. The scenario plays normally; only the voice output is absent.
Timeout errors
Section titled “Timeout errors”A slow turn is not a failed turn. A local model on CPU-only hardware can spend a minute or more on a single reply — mostly reading the prompt back in — and the app waits for it. The clock and bar described in How long a turn will take are on screen for the whole wait; after five seconds the app adds that the NPC is taking longer than usual, and after thirty seconds that the reply has not been thrown away.
Two separate limits can still end a turn early:
| Limit | Value | What happens |
|---|---|---|
| Engine went quiet | 180 s with nothing sent (CONVSIM_LLAMA_CPP_CHAT_TIMEOUT) — this covers reading the prompt back in as well as the gaps between words | The turn fails with a timeout error. Nothing is recorded, so you can retype the same turn. |
| App gave up waiting | 10 min | From 5 minutes on, the app stops trusting the request and starts asking the session what it actually recorded, every 15 seconds. As soon as the reply has landed it is shown and play continues. If the request itself answers while that is going on — with the reply, or with a failure of its own — that answer is used straight away. Only if nothing has landed by 10 minutes does the turn fail with a timeout error. (A session started with transcript saving off has nothing to ask about, so the app simply keeps waiting on the request until then.) |
The session is not ended by either case — you can retry the same turn. The error message includes the same suggestions listed above.
Turning off features to improve performance
Section titled “Turning off features to improve performance”All of the following can be changed without restarting a session:
| Feature | How to disable | Performance impact |
|---|---|---|
| TTS | Uncheck “Enable TTS” in scenario setup | Removes Kokoro synthesis latency |
| VAD | Switch to push-to-talk in scenario setup | Removes silence-detection overhead |
| State meters | Uncheck “Show state meters” in setup | Minor |
| Transcript saving | Uncheck “Save transcript” in setup | Minor |
Runtime Settings (GPU layers, context length, threads, temperature) take effect on the next model load.