Primeiros passosRankingDecode calculatorModelosReportsHardwareAvaliações comparativasMarketplaceAluguéisProDocs da API
Idioma
Actual Computer — Every computer, one endpoint
Community field report

Qwen3.8-27B long-context test on dual RTX 3060 12GB

Qwen3.8-27B on two RTX 3060 12GB cards, llama.cpp 0.4.1-dev. Measured 129,000-token synthetic prompt at 497.8 prompt tok/s and 42.7 output tok/s mean, with exact recall of 3/3 canary facts in two runs. Configured context 131,072; ngram-mod + MTP3; batch 2048, ubatch 512.

Long-context prefill and retrieval test on two NVIDIA RTX 3060 12GB cards (24GB combined, PCIe, no NVLink). Model: Unsloth Qwen3.8-27B-UD-Q4_K_XL. Engine: llama.cpp 0.4.1-dev, commit b29c606e2.

Configuration

  • Configured context: 131,072 tokens; text-only, no mmproj
  • Tensor split: --split-mode tensor --tensor-split 1,1
  • KV cache: K/V q4_0; flash attention on; one parallel slot
  • Speculation: ngram-mod,draft-mtp, MTP max draft 3
  • Batch 2048, ubatch 512; GGML_CUDA_ALLREDUCE=internal
  • Approx. 11,141 MiB used per GPU

Long-prompt results

Prompts are deterministic synthetic operations notes. Each contains exact canary facts near the start, middle, and end; the model returns them in JSON. Prompt throughput is llama-server timings.prompt_per_second; output throughput is timings.predicted_per_second.

Actual prompt tokensPrompt tok/sOutput tok/sWall timeExact canary recall
4,072664.9577.5310.6 s3/3
16,360692.6162.5227.2 s3/3
32,743652.3557.0257.4 s3/3
65,510594.7952.81114.7 s3/3
98,280537.5045.09188.6 s3/3
122,855503.2843.85251.9 s3/3
129,000497.83 mean (2 runs)42.70 mean264.6 s mean3/3 on both

The 129k prompt is 98.42% of the configured context; both runs produced the same prompt and answer hashes. From 16,360 to 129,000 input tokens, prompt-eval throughput declined 28.1%. All tested prompts completed and returned all three canaries.

Tuning comparison at 65,510 prompt tokens

ProfilePrompt tok/sOutput tok/sVRAM per GPU
ngram-mod + MTP3, ubatch 512594.7952.81~11,141 MiB
MTP3 only, ubatch 512593.4751.27~11,135 MiB
ngram-mod + MTP3, ubatch 256558.8352.02~10,921 MiB

MTP-only prefill was effectively tied. Ubatch 256 was about 6% slower in prefill while saving around 220 MiB per GPU. The measured winner for this setup is ngram-mod + MTP3 with ubatch 512.

Method and limits

The harness uses cache_prompt=false, temperature 0, a fixed seed, JSON output, and a discarded 12k-token warmup before the sweep. It samples llama-server metrics to detect overlapping requests and marks a run invalid if another request overlaps. The near-limit prompt was repeated twice; other sizes and the 65k profile variants were measured once. This synthetic exact-retrieval task provides a repeatable context/throughput check, not a broad real-world quality evaluation. Output speed is with speculative decoding enabled and should not be presented as a no-spec llama-bench result.

Reproduction harness and full machine-readable result: GitHub recipe repo, benchmarks/long_context_recall.py and results/2026-09-26-131k-context.json.

Discussion

0 comments

Questions, reproduction notes, and follow-up results.

No comments yet. Start the technical discussion.