Scientific Benchmark Comparison
Isolate independent variables across local LLM inference and agent workloads. NeuralDrift never collapses multi-dimensional performance into a single misleading composite score.
RUN A (BASELINE)
RUN B (CANDIDATE)
COMPARISON BLOCKED / SCIENTIFICALLY INVALID
These runs cannot be directly compared.
- Different workloads: ND-AGENT-001 vs ND-LLM-001.
- Different benchmark paradigms: workload vs raw_inference.
Direct Metric Comparison
| Metric | Run A: qwen2.5.1-coder-7b-instruct | Run B: qwen2.5.1-coder-7b-instruct | Observation |
|---|---|---|---|
| Workload | ND-AGENT-001 v1.0.0 | ND-LLM-001 v1.0.0 | Different Tasks |
| Hardware GPU | NVIDIA GeForce RTX 5080 | NVIDIA GeForce RTX 5080 | Same GPU |
| Runtime | lm_studio (0.3.9) | lm_studio (0.3.9) | Same Runtime |
| Outcome | PASS | PASS | Both PASS |
| Peak VRAM | 10.57 GB | 10.57 GB | Hardware footprint |