Guide: How to Optimize Performance on Any GPU
Complete optimization guide by VRAM tier
#GPU Tier Optimization
Low VRAM (4-6GB): GTX 1660, RTX 3050, RTX 4050
Settings:
Recommendations:
- →Use SD 1.5 ONLY
- →Max resolution: 512×512
- →Batch size: 1
- →Install Tiled VAE custom node
- →Steps: 20-25 max
- →Avoid SDXL completely
Medium VRAM (8-12GB): RTX 3060, RTX 4060
Settings:
Recommendations:
- →SDXL supported at 768×768
- →SD 1.5 at 768×768 or 1024×1024
- →Batch size: 1-2
- →Steps: 25-30
- →Use fp16 models when available
High VRAM (16-24GB): RTX 4080, RTX 4090
Settings:
Recommendations:
- →SDXL at 1024×1024 or higher
- →Batch size: 2-4
- →Steps: 30-50
- →High-res upscaling supported
- →Video generation supported
- →Multiple model loading
#Speed Optimization Tips
- →
Install xformers
comfyui-workflow.jsonThis can change memory use or speed for some installations; test one workflow before relying on it in production.
- →
Use NVIDIA GPU scheduling (Windows)
- →NVIDIA Control Panel → Manage 3D Settings
- →Low Latency Mode: Ultra
- →Power Management: Prefer Maximum Performance
- →
Close background apps during generation
- →
Use SSD for model storage (not HDD)
- →
Keep drivers updated
#Related Guides
#A repeatable optimization loop
Avoid changing several variables at once. Start with a workflow that already completes, record its model version, resolution, batch size, sampler, precision, and any launch flags. Change one setting, run the same input again, and keep the output only if it improves the result you care about without introducing an error.
For memory failures, lower batch size first, then reduce resolution or disable optional nodes. For slow first runs, distinguish model-loading time from generation time. For a workflow that will be reused, keep a short execution note with the exact graph and software versions. NeuralDrift labels compatibility and timing as unknown unless such a record exists.
#When to stop tuning
Do not keep increasing resolution, steps, or concurrent jobs just because a previous run worked. Stop when the output meets the task requirement, memory headroom is stable, and the run can be reproduced. If a change produces a new error, revert that one setting and capture the full error text before changing dependencies.