Long-context prefill and retrieval test on two NVIDIA RTX 3060 12GB cards (24GB combined, PCIe, no NVLink). Model: Unsloth Qwen3.8-27B-UD-Q4_K_XL. Engine: llama.cpp 0.4.1-dev, commit b29c606e2.
Configuration
- Configured context: 131,072 tokens; text-only, no mmproj
- Tensor split:
--split-mode tensor --tensor-split 1,1 - KV cache: K/V
q4_0; flash attention on; one parallel slot - Speculation:
ngram-mod,draft-mtp, MTP max draft 3 - Batch 2048, ubatch 512;
GGML_CUDA_ALLREDUCE=internal - Approx. 11,141 MiB used per GPU
Long-prompt results
Prompts are deterministic synthetic operations notes. Each contains exact canary facts near the start, middle, and end; the model returns them in JSON. Prompt throughput is llama-server timings.prompt_per_second; output throughput is timings.predicted_per_second.
| Actual prompt tokens | Prompt tok/s | Output tok/s | Wall time | Exact canary recall |
|---|---|---|---|---|
| 4,072 | 664.95 | 77.53 | 10.6 s | 3/3 |
| 16,360 | 692.61 | 62.52 | 27.2 s | 3/3 |
| 32,743 | 652.35 | 57.02 | 57.4 s | 3/3 |
| 65,510 | 594.79 | 52.81 | 114.7 s | 3/3 |
| 98,280 | 537.50 | 45.09 | 188.6 s | 3/3 |
| 122,855 | 503.28 | 43.85 | 251.9 s | 3/3 |
| 129,000 | 497.83 mean (2 runs) | 42.70 mean | 264.6 s mean | 3/3 on both |
The 129k prompt is 98.42% of the configured context; both runs produced the same prompt and answer hashes. From 16,360 to 129,000 input tokens, prompt-eval throughput declined 28.1%. All tested prompts completed and returned all three canaries.
Tuning comparison at 65,510 prompt tokens
| Profile | Prompt tok/s | Output tok/s | VRAM per GPU |
|---|---|---|---|
| ngram-mod + MTP3, ubatch 512 | 594.79 | 52.81 | ~11,141 MiB |
| MTP3 only, ubatch 512 | 593.47 | 51.27 | ~11,135 MiB |
| ngram-mod + MTP3, ubatch 256 | 558.83 | 52.02 | ~10,921 MiB |
MTP-only prefill was effectively tied. Ubatch 256 was about 6% slower in prefill while saving around 220 MiB per GPU. The measured winner for this setup is ngram-mod + MTP3 with ubatch 512.
Method and limits
The harness uses cache_prompt=false, temperature 0, a fixed seed, JSON output, and a discarded 12k-token warmup before the sweep. It samples llama-server metrics to detect overlapping requests and marks a run invalid if another request overlaps. The near-limit prompt was repeated twice; other sizes and the 65k profile variants were measured once. This synthetic exact-retrieval task provides a repeatable context/throughput check, not a broad real-world quality evaluation. Output speed is with speculative decoding enabled and should not be presented as a no-spec llama-bench result.
Reproduction harness and full machine-readable result: GitHub recipe repo, benchmarks/long_context_recall.py and results/2026-09-26-131k-context.json.
