qwen2.5.1-coder-7b-instruct
PASSNEURALDRIFT BENCHMARKEDWorkload ND-LLM-001 evaluated on NVIDIA GeForce RTX 5080 via lm_studio (0.3.9) using CUDA backend.
Measured Telemetry
Raw inference metrics capturing model execution throughput and sub-second token latency.
TIME TO FIRST TOKEN (TTFT)
66.8 ms
Streaming latency to initial token emission.
GENERATION THROUGHPUT
93.1 tok/s
Sustained token generation rate.
PROMPT PROCESSING
2140 tok/s
Context ingestion rate (143 tokens in 66.8 ms).
TOTAL TIME
1.02 s
Total wall-clock duration for 232 tokens.
PEAK VRAM OBSERVED
10.57 GB
Dedicated GPU memory usage during generation.
GPU UTILIZATION
2%
Active compute engine utilization.
Objective Validation
VALIDATOR: JSON_SCHEMA
VALIDATION PASSEDAll required JSON keys and values matched.
OUTPUT RECEIVED
{
"parameters_b": 14800000000,
"bytes_per_param": 2,
"weights_gb": 29600.0,
"kv_overhead_gb": 1.25,
"total_vram_gb": 29711.25,
"fits_16gb_gpu": false
}Environment & Reproducibility
HARDWARE PROFILE
| GPU | NVIDIA GeForce RTX 5080 |
| Architecture | Blackwell |
| VRAM Capacity | 15.92 GB |
| System RAM | 61.64 GB |
| CPU | AMD Ryzen 9 9950X3D 16-Core Processor |
| Form Factor | DESKTOP |
SOFTWARE & RUNTIME
| Runtime | lm_studio (0.3.9) |
| Backend | CUDA 13.4 |
| GPU Driver | 616.92 |
| Operating System | win32 10.0.26200 |
| Model Quantization | GGUF / Q4_K_M |
| Context Limit | 32768 tokens |
CONFIGURATION FINGERPRINT
Deterministic SHA-256 hash across model, quantization, runtime, runtime version, GPU, driver, workload version, and generation settings. Two runs must share material variables to be directly comparable.
108fb295b3130c072e3126b845837ae53cd1638d7960a88a7299c96232d1f552Run ID: run-2026-09-17T13-27-58-678Z-nd-llm-001-qwen2_5_1-coder-7b-instruct-rtx-5080-c1fead83 · Recorded at: 2026-09-17T13:27:59.702Z · Benchmark Suite v1.0.0