← NeuralDrift Lab/COMPARISON ENGINE

Scientific Benchmark Comparison

Isolate independent variables across local LLM inference and agent workloads. NeuralDrift never collapses multi-dimensional performance into a single misleading composite score.

RUN A (BASELINE)

RUN B (CANDIDATE)

COMPARISON BLOCKED / SCIENTIFICALLY INVALID

These runs cannot be directly compared.

  • Different workloads: ND-AGENT-001 vs ND-LLM-001.
  • Different benchmark paradigms: workload vs raw_inference.

Direct Metric Comparison

MetricRun A: qwen2.5.1-coder-7b-instructRun B: qwen2.5.1-coder-7b-instructObservation
WorkloadND-AGENT-001 v1.0.0ND-LLM-001 v1.0.0Different Tasks
Hardware GPUNVIDIA GeForce RTX 5080NVIDIA GeForce RTX 5080Same GPU
Runtimelm_studio (0.3.9)lm_studio (0.3.9)Same Runtime
OutcomePASSPASSBoth PASS
Peak VRAM10.57 GB10.57 GBHardware footprint