← NeuralDrift Lab/BENCHMARK RESULT

qwen2.5.1-coder-7b-instruct

PASSNEURALDRIFT BENCHMARKED

Workload ND-LLM-001 evaluated on NVIDIA GeForce RTX 5080 via lm_studio (0.3.9) using CUDA backend.

Measured Telemetry

Raw inference metrics capturing model execution throughput and sub-second token latency.

TIME TO FIRST TOKEN (TTFT)

66.8 ms

Streaming latency to initial token emission.

GENERATION THROUGHPUT

93.1 tok/s

Sustained token generation rate.

PROMPT PROCESSING

2140 tok/s

Context ingestion rate (143 tokens in 66.8 ms).

TOTAL TIME

1.02 s

Total wall-clock duration for 232 tokens.

PEAK VRAM OBSERVED

10.57 GB

Dedicated GPU memory usage during generation.

GPU UTILIZATION

2%

Active compute engine utilization.

Objective Validation

VALIDATOR: JSON_SCHEMA

VALIDATION PASSED

All required JSON keys and values matched.

OUTPUT RECEIVED

{
  "parameters_b": 14800000000,
  "bytes_per_param": 2,
  "weights_gb": 29600.0,
  "kv_overhead_gb": 1.25,
  "total_vram_gb": 29711.25,
  "fits_16gb_gpu": false
}

Environment & Reproducibility

HARDWARE PROFILE

GPUNVIDIA GeForce RTX 5080
ArchitectureBlackwell
VRAM Capacity15.92 GB
System RAM61.64 GB
CPUAMD Ryzen 9 9950X3D 16-Core Processor
Form FactorDESKTOP

SOFTWARE & RUNTIME

Runtimelm_studio (0.3.9)
BackendCUDA 13.4
GPU Driver616.92
Operating Systemwin32 10.0.26200
Model QuantizationGGUF / Q4_K_M
Context Limit32768 tokens

CONFIGURATION FINGERPRINT

Deterministic SHA-256 hash across model, quantization, runtime, runtime version, GPU, driver, workload version, and generation settings. Two runs must share material variables to be directly comparable.

108fb295b3130c072e3126b845837ae53cd1638d7960a88a7299c96232d1f552

Run ID: run-2026-09-17T13-27-58-678Z-nd-llm-001-qwen2_5_1-coder-7b-instruct-rtx-5080-c1fead83 · Recorded at: 2026-09-17T13:27:59.702Z · Benchmark Suite v1.0.0