Running PhoneLLM Alpha 1 on RTX 3090: VRAM, Latency, Callers
Waffles is an AI voice agent living on our softcery.com website. She runs on Gemma 4 26B-A4B, hosted on a rented RTX 3090, the card and the flags in lab 038.
Then Daily shipped PhoneLLM Alpha 1, a full-parameter fine-tune of Nemotron 3 Nano 30B-A3B, built for phone agents. On PhoneBench Alpha 1 it scores 72.3, against 72.4 for GPT-5.6 Terra, 78.6 for Gemini 3.6 Flash and 58.1 for Gemma 4 31B. That page publishes no judge, no item list and no harness, and its Gemma row states no quant and no thinking setting.
Daily serves it on a B200 in BF16, 44 processes per card, and states 331 ms p50 to the first token and $0.0025 per minute at 70 % utilization.
So: does it serve the same card at the same turn budget? This article covers serve and load. Quality goes in a second article, once we have a scorer.
Two hosts carry the rows: an RTX 3090 on vast.ai in Czechia at $0.260 per hour, destroyed after its runs, and an RTX 3090 on vast.ai in Spain at $0.188. A row is from the Spain host unless it says Czechia.
The outcome
On a rented RTX 3090 on vast.ai at $0.19 to $0.26 per hour ($140 to $190/mo + network) + mradermacher/phonellm-alpha-1-i1-GGUF, the IQ4_XS file, llama.cpp b10902:
llama-server \
-m phonellm-alpha-1.i1-IQ4_XS.gguf \
-ngl 99 -c 393216 -np 12 -ub 2048 --cache-type-k f16 --cache-type-v f16 -fa on \
--no-cache-idle-slots \
--jinja --chat-template-kwargs '{"enable_thinking": false}' \
--temp 0 --top-k 40 --top-p 1.0 --min-p 0 --repeat-penalty 1.0
-np 12 gives 12 slots of 32k. -np 24 at the same -c gives 24 slots of 16k, twice the callers, section 6.
| Metric | Result | Context |
|---|---|---|
| VRAM | 20.7 GiB of 24 | 12 slots of 32k, 240 MiB each, 3.1 GiB free. 24 slots of 16k leave 2.6 GiB |
| Warm turn, 4 new tokens | 67 to 120 ms to the first token | 4 to 54 messages, with and without 5 tool schemas, 3 reps each |
| Warm turn, replay | 188 to 220 ms to the first token | Czechia. 10.4k-token history, 24 to 62 new tokens, 3 reps of 20 to 24 cases. The A4B: 91 to 109 ms |
| Prompt eval floor | 47.6 ms | 4 new tokens. 35.1 ms of it is checkpoint work, section 2. 54 ms on Czechia. The A4B: 10 to 13 ms |
| Decode, one caller | 213.6 to 215.7 t/s | 15 one-caller rows over 5 hours, card 62 to 68 C |
| Decode, 12 callers | 47.6 t/s each | all 12 talking at once, 550 ms p50 to the first token, 0.221 of the one-caller rate |
| Decode, 24 callers | 28.4 t/s each | 24 slots of 16k, 1,083 ms p50, 682 t/s in total |
| New call, card idle | 6,035 t/s prompt eval | 17,878 tokens in 2,962 ms. The A4B: 4,200 t/s |
| New call, 11 warm callers | 2,470 to 2,858 ms to the first token | prompt eval 2,174 to 2,366 ms. The 11 warm callers wait 1,862 ms p50, 2,230 max |
| Unique prompts | 1,587 ms p50 at 1 request, 5,173 at 4 | guidellm, 10,016-token prompts, 130 requests. 5,025 to 6,634 t/s input, 0.50 to 0.66 new calls per second |
| Barge-in reach | 1,200 to 1,600 tokens | the deepest edit that avoids a full recompute. -ub 2048, 394 to 449 ms inside it. 320 to 360 tokens at -ub 512 |
| Empty replies | 0 in 396 | replay and conversation runs at 3 and 12 slots, Czechia |
| Tool calls | 9 of 9 as tool_calls |
3 tool cases, 3 reps each. 27 fire on the 66-case replay. The Nemotron format, parsed by the server |
| Cost | $0.00013 per voice minute | 24 callers at 100 % occupancy on $0.188 per hour, by division. $0.00026 at 12 callers. $0.00024 and $0.00048 at 55 % |
Daily's 331 ms p50 is a first token under load with unique prompts. Compare it with our unique-prompt row, 1,587 to 5,173 ms.
1. Layers, slot cost and the quant
PhoneLLM is a Nemotron 3 Nano hybrid: 52 layers, of which 23 are Mamba2 state layers, 23 are Mixture of Experts, and 6 are attention. Only the 6 attention layers keep a KV cache, one entry per token. The 23 Mamba2 layers keep a state each, a running summary of the history. The 23 states cost 48 MiB per slot to hold. A rewind needs a checkpoint, section 3.
The cost of one caller: 6 attention layers with 2 KV heads give 6 KiB per token, 192 MiB for a 32k slot. The 23 Mamba2 states add 48 MiB per slot in f32. A slot is 240 MiB at 32k and 144 MiB at 16k. The A4B slot is 940 MiB at 32k.
The weights are the second limit. The expert rows are 2,688 and 1,856 wide. Neither divides by 256, so each 256-block quant type falls back to a 32-block type on the 46 expert tensors. For 7 of the 9 public files the experts are IQ4_NL or Q4_0 and weigh 16.5 GB. We read the GGUF headers of the 9 public files over HTTP range requests, no download:
| Public file | Expert type | Size | Free at 12 slots of 32k |
|---|---|---|---|
| IQ1_S to IQ4_XS | IQ4_NL or Q4_0 | 17.9 to 18.2 GB | 3.1 to 3.8 GiB, measured on 2 hosts |
| Q4_K_S | Q5_0 | 21.9 GB | under 0.3 GiB by the measured overhead |
| Q4_K_M | Q5_0 | 24.5 GB | none |
The 7 files differ by 0.3 GB. The quantizer's model card does not say so; it is one template for every model. We took IQ4_XS, the top of the 18 GB group, and there is no smaller quant to try on this card.
Side by side on a 24 GB card:
| 26B-A4B | PhoneLLM | |
|---|---|---|
| Weights | 14.25 GB | 17.97 GB |
| Slot at 32k | 940 MiB | 240 MiB |
| Free VRAM | 0.8 GiB at 10 slots | 3.1 GiB at 12 slots, 2.6 GiB at 24 slots of 16k |
PhoneLLM needs 3.7 GB more for weights and 8.2 GiB less for 12 slots.
2. The 47.6 ms floor
A warm turn has 3 parts. The server renders the chat template, it does its other work before the response headers, and it runs the prompt eval over the new tokens. POST /apply-template gives the render alone, with no slot and no prompt eval. timings.prompt_ms gives the prompt eval. The time before the headers includes the render. The turn is that time plus the prompt eval.
One caller, a cached history of the recorded call, and a prompt that adds 4 new tokens, 3 reps each:
| Messages | Tools | Render | Before headers | Prompt eval | Turn |
|---|---|---|---|---|---|
| 4 | no | 3 ms | 20 to 22 ms | 47 to 49 ms | 67 to 71 ms |
| 4 | 5 | 14 ms | 41 to 46 ms | 48 ms | 89 to 94 ms |
| 22 | no | 8 ms | 29 to 31 ms | 47 to 48 ms | 76 to 79 ms |
| 22 | 5 | 19 ms | 47 to 51 ms | 47 to 48 ms | 94 to 99 ms |
| 54 | no | 17 to 18 ms | 43 to 47 ms | 47 to 48 ms | 91 to 94 ms |
| 54 | 5 | 29 to 30 ms | 67 to 72 ms | 48 ms | 115 to 120 ms |
The prompt eval is 46.9 to 49.2 ms on all 18 rows, 47.6 p50. It holds at 4 and 54 messages, with and without the 5 tool schemas. It holds at -ub 512, 2048 and 4096, at checkpoints 2, 8, 32 and the default, and with flash attention off. The Czechia host reads 54 to 55 ms on the same test. The A4B runs a 1-token prompt eval in 10 to 13 ms.
The server log at verbosity 4 splits one 46.91 ms prompt eval:
| Step | Cost |
|---|---|
| Checkpoint restore | 8.05 ms |
| Sampler init | 1.21 ms |
| Checkpoint create | 27.05 ms |
| The 4-token prompt eval | 10.58 ms |
Checkpoint work is 35.10 ms of 46.91, 75 %. On each warm turn the server writes a 47.618 MiB checkpoint at token 10,369, the end of the history, and it supersedes the checkpoint it wrote there on the turn before. One create runs at 1.76 GB per second. No flag in b10902 skips a create at a token that holds one. That is an upstream change, section 10.
The prompt eval grows with the new tokens, less than linearly. On the Czechia replay of the recorded call, 74 warm turns of 26 to 105 new tokens, the prompt eval is 135.6 ms p50 at 60 tokens. On the conversation run, the recorded caller turns sent live against the model's own replies, it is 130 ms at 24 tokens. So 54 ms at 4 tokens, 130 at 24, 136 at 60, on that host. The step between 4 and 24 is unmeasured. The A4B is at 40 ms for 61 tokens.
The render is CPU work and follows the host. On Spain it costs 3 ms at 4 messages, 8 at 22 and 18 at 54, and the 5 tool schemas add 11 to 12 ms. On Czechia it cost 5, 18 and 35 to 43 ms, and the schemas added 25 to 32. The Gemma template renders in 1.6 ms at 82 messages. PhoneLLM's total before the headers is still lower: 67 to 72 ms at 54 messages with tools, where the A4B spent 126 to 139 ms at 42 messages. The A4B's 2.5 ms per history message does not appear on PhoneLLM.
So the 100 ms gap in the replay warm turn, 188 to 220 ms against 91 to 109, is the prompt eval: the checkpoint floor plus the unmeasured step to 24 tokens. The render and the headers add nothing to the gap.
3. Checkpoints
On the 12B, --ctx-checkpoints 0 cut the warm turn from 556 ms to 72. That server took 2 or 3 checkpoints of 320 MiB on each request, and a warm turn did not need them. We carried the flag over to PhoneLLM and ran the same header test as in section 2 on the Czechia host. All 12 rows came back with cache_n 0:
| Messages | Tools | Prompt tokens | Prompt eval, 3 reps |
|---|---|---|---|
| 4 | no | 7,286 | 2,130 to 4,311 ms |
| 4 | 5 | 9,992 | 3,629 to 5,853 ms |
| 54 | no | 9,928 | 4,213 to 6,055 ms |
| 54 | 5 | 12,634 | 5,977 to 7,727 ms |
A repeat of an identical 17,878-token prompt gave the same result: cache_n 0, 8,262 ms.
The cause is the recurrent state. A KV cache holds one entry per token, so the server can cut it at any position and continue from there. A Mamba2 state is a fixed-size summary of the whole history up to a position. To continue from an earlier position, the server needs a copy of the state saved at that position. A checkpoint is that copy. With 0 checkpoints there is no copy, and the server starts from token 0 on each turn, including a turn whose prompt matches the cache in full.
The trace shows the position of each copy. The server creates a checkpoint at the end of a prompt eval chunk. One cold 17,878-token prompt at -ub 2048 creates 4, at tokens 9,964, 10,366, 15,829 and 17,873, 47.618 MiB each. On a rewind the server tests each checkpoint newest first against the common prefix, restores the newest valid one and recomputes the tail. At 400 tokens back it rejects 17,873, restores 15,829 and recomputes 2,248 tokens in 419 ms. At 1,600 back it restores 10,366 and recomputes 8,311 tokens in 1,335 ms.
The count binds at no measured value. The server creates 4 checkpoints on a cold prompt and 5 after a rewind, under each of the 4 caps from 2 to 32:
--ctx-checkpoints 2, 8 and 32. Each gives the same rewind rows as the default, token for token. The floor reads 47.4 to 47.7 ms on all of them.--checkpoint-min-step 0and 2048, against the 8192 default. Same rows. The spacing does not set the cut.--cache-ram 0and -1, against the 8192 MiB default. Same rows. That budget holds the RAM prompt cache. The checkpoints sit outside it.
So the count does not move the cost, and the ubatch does, section 4. The defaults stay. A flag that saved 480 ms per turn on Gemma costs 2 to 8 s per turn on a hybrid.
Two upstream issues touch this path. Issue 22384 reports that checkpoint creation needs 64 tokens or more, and issue 24055 reports that checkpoints are invalidated on hybrid models from build b9354. On b10902 the trace does create and restore checkpoints, and the warm turn reuses the cache at 47.6 ms, so neither defect holds here.
4. Barge-in
When the caller interrupts, the agent cuts the reply and sends the next turn with the cut reply in the history. To the server that is a prompt whose prefix ends before the end of the cache. On a KV cache the server drops the tail and continues. On this model it restores the nearest checkpoint before the cut and recomputes from there, section 3.
We cached a 17,878-token prompt, the 10,367-token recorded call plus a 7,511-token filler tail, then replaced the last N tokens of the tail and read cache_n, prompt_n and prompt_ms on the next turn. Czechia rows at -ub 256, 512 and 1024, one rep per point. Spain rows at 2048 and 4096:
-b, -ub |
Reach, tokens back | Inside the reach, tokens recomputed | Inside, prompt eval | Past the reach |
|---|---|---|---|---|
| 2048, 256 | 50 to 200 | 286 | 166 ms | 7,611 tokens, 2,579 ms |
| 2048, 512 | 320 to 360 | 676 | 179 to 182 ms | 7,691 tokens, 1,552 to 1,765 ms |
| 2048, 1024 | 400 to 800 | 1,202 | 315 ms | 7,911 tokens, 1,955 ms |
| 2048, 2048 | 1,200 to 1,600 | 2,098 to 2,448 | 394 to 449 ms | 8,311 to 9,011 tokens, 1,334 to 1,425 ms |
| 4096, 4096 | 2,000 to 2,800 | 4,296 to 4,496 | 722 to 756 ms | 8,911 to 10,011 tokens, 1,431 to 1,570 ms |
Inside the reach the server recomputes about one ubatch plus the edit. Past it, cache_n drops to 10,367 on each past-reach row in the table: the server restores the checkpoint at the end of the recorded call and recomputes the whole tail.
The default -b is 2048 and it caps -ub. -ub 4096 alone gives the same checkpoint positions as -ub 2048, 15,830 and 16,030. With -b 4096 the checkpoints move to 13,782, 13,982 and 14,382, one 4096 ubatch back from the prompt end, and the reach doubles. That setting costs 542 MiB of VRAM and moves the floor 0.3 ms.
The trace names the reach. It is the distance from the prompt end to the newest checkpoint the server accepts, and the checkpoints sit at prompt eval chunk ends. One gap stays. At 1,600 back on -ub 2048 the server rejected the checkpoint at 16,030, which sat 248 tokens before the edit, and fell to 10,367. The margin of that check is not read from the source. The Gemma article has a gap of the same kind, 400 against 512.
On the A4B the reach is set by the 1,024-token sliding window and does not move with the batch. On PhoneLLM the ubatch sets it. A spoken reply is 6 to 107 tokens on our calls, so a barge-in cut stays inside the reach on both models, at any ubatch from 512 up. The larger ubatch gives a deeper edit at a higher cost: at 2048 a change 1,200 tokens back costs 394 to 449 ms, at 4096 it costs 722 to 756. A trimmed tool result or a rewritten system line inside 1.2k tokens stays under 0.45 s at 2048.
-ub 256 is rejected: the reach falls under 200 tokens and a deep edit costs 2.6 s. -b 4096 -ub 4096 is measured and not adopted: it halves the cold join stall, section 5, and doubles the cost of each barge-in. We run 2048.
5. The second caller
A new caller arrives with a prompt that shares nothing with the slots: a cold 10,377-token prompt eval. The server runs all active slots in one batch, so the cold prompt goes through in chunks of one ubatch, and each decode step of a warm caller runs in the batch with the next chunk. The warm caller gets no token until the cold prompt is through.
One cold caller beside 11 warm callers on the 12-slot config, all 12 requests fired at once, 3 reps, card 67 C:
-b, -ub |
Cold caller, first token | Cold prompt eval | Warm callers, first token p50, max | Warm rows over 500 ms |
|---|---|---|---|---|
| 2048, 2048 | 2,470 to 2,858 ms | 2,174 to 2,366 ms, 4.4 to 4.8k t/s | 1,862 ms, 2,230 ms | 27 of 33 |
| 4096, 4096 | 2,289 to 2,680 ms | 1,958 to 2,178 ms | 1,418 ms, 1,776 ms | 18 of 22 |
The stall has 2 shapes. In 27 of 33 warm rows the caller waits for the first token. In the other 6 the first token comes at 105 to 198 ms and the reply then decodes at 12 to 14 t/s, 3 times the 4.5 t/s of speech. The Gemma article saw both shapes, with the stall after the first word in 2 of 3.
The Czechia host ran the same join beside 2 warm callers. The cold prompt eval took 3,070 to 4,245 ms at -ub 2048 and 6,816 to 8,712 ms at the default 512. The warm callers stalled 2,229 to 3,245 ms and 6,267 to 7,888 ms. The A4B on the same join, at its 2,048-token chunk, stalls 2.2 to 2.8 s. At the same chunk size the two models match.
Prompt eval alone, one caller, card idle: 17,878 tokens in 2,962 ms, 6,035 t/s. The Czechia host gave 3,594 t/s on a 10,394-token prompt at an unknown temperature. The A4B does 4,200 t/s.
Not run: -b below the ubatch, the fix the Gemma article proposes for the stall. On this model the ubatch also sets the barge-in reach, section 4.
6. Callers per card
An idle slot costs 240 MiB of VRAM at 32k, 144 MiB at 16k, and no decode time. The cost comes with the active callers, because each decode step runs the whole batch. Each load point has a one-caller row measured right before it, and the column to read is the share: the decode per caller over that one-caller rate.
All slots warm on one cached 10,369-token prompt. All callers talk at once, replies of 40 to 51 tokens, 2 reps, card 62 to 68 C.
12 slots of 32k, one-caller rows at 214.5 t/s:
| Callers talking at once | First token, p50 | First token, max | Decode per caller, p50 | Share of one caller |
|---|---|---|---|---|
| 1 | 82 ms | 90 ms | 214.5 t/s | 1 |
| 4 | 172 ms | 224 ms | 94.7 t/s | 0.442 |
| 8 | 394 ms | 399 ms | 59.1 t/s | 0.276 |
| 12 | 550 ms | 559 ms | 47.6 t/s | 0.222 |
24 slots of 16k, one-caller rows at 213.6 to 215.4 t/s:
| Callers talking at once | First token, p50 | First token, max | Decode per caller, p50 | Share of one caller | Total decode |
|---|---|---|---|---|---|
| 4 | 216 ms | 224 ms | 102.8 t/s | 0.479 | 411 t/s |
| 8 | 376 ms | 395 ms | 59.5 t/s | 0.277 | 476 t/s |
| 12 | 559 ms | 567 ms | 47.4 t/s | 0.221 | 569 t/s |
| 16 | 739 ms | 751 ms | 40.8 t/s | 0.191 | 653 t/s |
| 20 | 896 ms | 919 ms | 32.5 t/s | 0.152 | 650 t/s |
| 24 | 1,083 ms | 1,104 ms | 28.4 t/s | 0.133 | 682 t/s |
The slot length does not set the rate. 12 slots of 32k and 24 slots of 16k give 59.1 against 59.5 t/s at 8 callers, and 47.6 against 47.4 at 12. The slot count is the ceiling. The slowest of the 48 rows at 24 callers decodes 18.5 t/s, 4 times the speech rate.
The first token grows 43 to 47 ms per added caller on both configs. The A4B grew 16 ms per caller, 149 ms at 1 to 290 at 10. That slope sets the count. VRAM leaves room: 24 slots of 32k load with 1,032 MiB free on the Czechia host, and 24 slots of 16k leave 2,592 MiB. The A4B at 10 callers gives 48 t/s per caller at 290 ms. PhoneLLM at 12 gives 47.6 t/s at 550 ms, and at 24 gives 28.4 t/s at 1,083 ms.
The Mamba2 state is 48 MiB at either length, so half the context cuts the slot by 40 %. 24 slots cost 3,456 MiB against 2,880 for 12. The longest prompt over 1,204 recorded call rows is 12,634 tokens, which a 16k slot holds with 3,750 spare. Our website agent reaches 24,568 tokens on one call, which a 16k slot truncates. Use 12 slots of 32k where a call can pass 16k. Use 24 slots of 16k where it cannot.
The rows above hold one cached prompt per slot. The service sees unique prompts. guidellm 0.7.3 at 1, 4, 12 and 24 concurrent requests, 10,016-token synthetic prompts, 48 output tokens. 130 requests, all successful, on 24 slots of 16k, card 41 to 73 C:
| Concurrent requests | First token, p50 | First token, p95 | Request, p50 | Input, t/s |
|---|---|---|---|---|
| 1 | 1,587 ms | 1,589 ms | 1.8 s | 6,634 |
| 4 | 5,173 ms | 7,113 ms | 7.5 s | 5,979 |
| 12 | 5,029 ms | 20,269 ms | 22.5 s | 5,155 |
| 24 | 3,499 ms | 137,481 ms | 23.2 s | 5,025 |
The card evaluates 5,025 to 6,634 prompt tokens per second in total, at each of the 4 request counts. At 10,016 tokens per call that is 0.50 to 0.66 new calls per second. Prompt eval binds the service. The p95 of 137 s at 24 requests is a queue wait. On unique prompts the first token is 3 to 5 times the slot row on the same card.
The first version of our load script warmed one slot before the run. Each other caller paid a cold 10k prompt eval inside its measured turn, and the first-token maxima read 25 to 74 s. Warm each slot before a load run.
7. Card state
The Czechia host fell 55 % inside one run. One 24-minute sweep, same server, same config, a one-caller row of 43 to 46 tokens at the start and at the end:
| Point | Card temperature | Clocks | Power | Throttle reason | Decode, one caller |
|---|---|---|---|---|---|
| Sweep start | 52 C | 1,905 MHz | 338 W | none | 218.6 t/s |
| Sweep end | 86 C | 1,695 of 2,100 MHz | 166 of 350 W | none | 97.6 t/s |
The Spain host keeps its rate. Its 15 one-caller reference rows over 5 hours read 213.6 to 215.7 t/s at 62 to 68 C, a spread of 1 %. Its card reports a reason where the Czechia card reported none: SwPowerCap at 1 caller, 288 to 312 W at 1,740 to 1,935 MHz, and SwThermalSlowdown at 67 C under 12 callers.
A card at 1,695 MHz with no reason and a card at 1,740 MHz with SwPowerCap differ. The Czechia card drew 166 W and set no flag, so its limit sat outside what nvidia-smi reports. 1,695 MHz is the rated boost clock of a 3090, 1.70 GHz on the NVIDIA spec page: the clock fell 11 % and the decode fell 55 %. The GDDR6X junction temperature is the candidate. Decode on this model is bound by memory bandwidth, and the GDDR6X on a 3090 runs 20 to 30 C above the GPU sensor and throttles on its own junction temperature, measured at up to 110 C on these cards. nvidia-smi does not report that sensor, an open request on the NVIDIA developer forum. It is not in the log on either host.
So the 2.2x fall is a defect of one host and does not reproduce on the second. The rule has 3 steps. Log the card, the clocks, the power and the throttle reasons on each measured row. Measure a one-caller row next to each load point. Publish the share. That is why section 6 carries shares and card temperatures.
On one boot on the Spain host the server ran without the card. ggml_cuda_init logged that it detects no CUDA-capable device, the health check answered ok, and the server decoded on CPU at 6.9 t/s. The card fields on those rows are null, so the rows are visible as invalid. It happened twice, both on a restart. A vast stop and start returned the same card by UUID both times. The health check does not test the card.
8. Renting
Two hosts, both a 1× RTX 3090 on vast.ai on the same image: Czechia at $0.260 per hour, destroyed after its runs, and Spain at $0.188. The weight pull is 17.97 GB. We did not time the boot on this build.
The image pins llama.cpp b10902, released 2026-09-11. The 3 upstream changes this model needs all merged before the pin: Nemotron 3 Nano support on 2025-12-16, the Mamba2 shape fix on 2026-06-26 and the checkpoint skip on 2026-06-11. MTP merged on 2026-05-16 as well, and the build does not use it. It carries no draft head: on a hybrid each draft step pays a checkpoint restore, so we did not test it.
The same replay gives 201 ms to the first token over loopback and 311 ms from a laptop: 110 ms per turn for the trip to Czechia. The Gemma host in Sweden cost 150 to 230 ms from the same laptop.
Cost per voice minute, the hourly price over the callers by division:
| Recipe | Card, price | Callers | First token, p50 | Decode per caller | $ per voice minute |
|---|---|---|---|---|---|
| This build, 24 slots of 16k | RTX 3090, $0.188/h | 24 | 1,083 ms, cached | 28.4 t/s | $0.00013 at 100 %, $0.00024 at 55 % |
| This build, 12 slots of 32k | RTX 3090, $0.188/h | 12 | 550 ms, cached | 47.6 t/s | $0.00026 at 100 %, $0.00048 at 55 % |
| phone-llm-fp8, vLLM on Modal, its README | L40S 48 GB, $1.95/h | 16 | 486 ms | 57 t/s | $0.00203 |
| Daily, BF16 on a B200, its blog | B200 | 44 | 331 ms, unique | not stated | $0.0025 at 70 % |
The 4 rows do not share one method. The FP8 row and the Daily row come from their authors and we re-measured neither. Our cached rows understate the first token 3 to 5 times against unique prompts, section 6.
9. What we did not measure
- Quality. The replays give counts: 0 empty replies in 396 rows, 27 tool calls in 198 rows. Two counts stand against the model: the arguments match the recorded ones in 3 of 9 calls, and the reply mirrors the caller's words in 4 of 18 cases, against 1 of 18 on the A4B. No scorer, no judge, and the suite prompts a Gemma persona. The second article covers quality.
- The step in prompt eval between 4 and 24 new tokens, 54 to 130 ms. No row between.
- The margin of the checkpoint check, the 248 tokens in section 4. Not read from the source.
- The cause of the hot Czechia card. Memory junction temperature and
clocks.memwere not logged. The Spain host does not reproduce it. - A mixed arrival on guidellm. The 4 rows hold a fixed request count, and the p95 at 24 requests shows a queue.
- A replay on 24 slots of 16k. The quality counts come from 3 and 12 slots of 32k.
- The boot time on this build.
- PhoneBench. Daily publishes no harness, item list or judge names.
- The BF16 baseline and the 24.5 GB Q4_K_M. Neither fits 24 GB.
- A phone-transcript imatrix. It needs a card over 24 GB.
--cache-reuse. Upstream disables it on each hybrid context, so no rental tests it.- A second prompt. Each number here is one recorded call of 10,369 to 10,377 tokens, 1 to 3 reps, on 2 hosts. The A4B rows come from the same call, 10,335 tokens in its tokenizer.
10. What to try next
In run order:
| Try | Targets | Expected | Cost |
|---|---|---|---|
| A checkpoint create that skips a token that holds one | 27.05 ms of the 46.91 ms floor | a floor near 20 ms. The server rewrites the same 47.6 MiB at token 10,369 on each warm turn | an upstream patch, one run |
| Prompt eval at 8, 12 and 16 new tokens | the step from 48 to 130 ms | the shape of the curve between the 2 measured points | one run |
| A shorter tool block | the 11 to 12 ms of render the 5 schemas add | render between 18 and 30 ms at 54 messages | agent change |
| guidellm with a Poisson arrival at 0.3 to 0.6 calls per second | the queue behind the 137 s p95 | a first-token p95 at the arrival rate the card holds | one run |
| The replay on 24 slots of 16k | the config that doubles the callers | the same counts as 12 slots of 32k | one run |
| Read the checkpoint validity check in the server source | the 248-token margin | the rule, and with it the exact reach per ubatch | reading |
Not on the list:
- A smaller quant. 7 of the 9 public files hold the same 16.5 GB of experts, section 1.
- The sampler. We run Daily's settings.
- An 8-bit KV cache. The cache is 192 MiB per slot and 2.6 to 3.1 GiB is free.
- Flash attention off. At 12 slots of 32k it asks for an 8.5 GiB compute buffer and the boot exits. At 1 slot it costs 2 GiB, 13 % of decode, and moves the floor 0.1 ms.
-b 4096 -ub 4096. Measured, sections 4 and 5. It doubles the barge-in cost.- More checkpoints, a smaller min step or a larger cache budget. All measured, section 3. None moves a row.