Stable Diffusion 1.5 VRAM Requirements
Stable Diffusion 1.5 is the low-VRAM baseline:
- 4GB VRAM: minimum with
--lowvram; expect slower generation and more compromises - 6GB VRAM: usable for many SD 1.5 512x512 workflows
- 8GB VRAM: comfortable for SD 1.5 and a better default if you do not want constant memory tuning
If your search is stable diffusion 1.5 minimum vram 4gb 6gb, the practical answer is: 4GB can run SD 1.5 with low-memory settings, 6GB is more usable, and 8GB is the safer recommendation. SDXL is a different tier and usually needs more headroom.
SDXL VRAM Requirements (April 2026)
SDXL requires 8GB VRAM minimum and 12GB recommended for stable 1024x1024 generation as of April 2026.
- Minimum: 8GB (with
--medvramoptimization) - Recommended: 12GB
- High-performance (ControlNet + LoRAs): 24GB
At 8GB, SDXL runs but requires memory optimization flags. At 12GB, SDXL runs natively at 1024x1024 without workarounds. At 24GB (RTX 4090), you can stack ControlNet, multiple LoRAs, and upscaling without hitting limits. If your local GPU falls short, cloud RTX 4090 cards start at $0.34/hr on RunPod.
Full VRAM Requirements by Model
| Task | Model | Min VRAM | Recommended | Example GPU |
|---|---|---|---|---|
| Generate 512x512 | SD 1.5 | 4 GB | 8 GB | RTX 3060 |
| Generate 1024x1024 | SDXL | 8 GB | 12–16 GB | RTX 4070 Ti Super |
| SDXL + ControlNet + LoRAs | SDXL | 12 GB | 16–24 GB | RTX 4080 / 4090 |
| Generate 1024x1024 | SD 3.5 Large | 12 GB (quantized) | 16–24 GB | RTX 4080 / 4090 |
| LoRA training | SDXL | 12 GB | 24 GB | RTX 3090 / 4090 |
| Full fine-tuning | Any | 24 GB | 48+ GB | A100 / H100 |
Minimum numbers assume optimization flags like --lowvram or --medvram. Recommended numbers are for comfortable generation without memory workarounds.
A card that handles SD 1.5 at 512x512 without issues will choke on SDXL at 1024x1024 with ControlNet loaded. This guide breaks down VRAM requirements by model, resolution, and task — so you size the hardware to the workload, not the other way around. If you have ever wondered why a faster GPU can actually feel slower, the answer is almost always that the workload spilled past the VRAM ceiling on the "faster" card.
Minimum GPU Requirements for Stable Diffusion
4–8 GB VRAM: SD 1.5 Only
Cards with 4–8 GB VRAM (GTX 1650, RTX 3050, RTX 4060) can run SD 1.5 at 512x512. At 4 GB, you need --lowvram mode, which offloads model layers to system RAM and slows generation significantly. At 8 GB, SD 1.5 runs natively without memory workarounds.
These cards cannot run SDXL at native resolution. Attempting SDXL on 8 GB VRAM without heavy optimization flags will produce out-of-memory errors. If you are stuck with 8 GB, see our CUDA out-of-memory error guide for workaround options.
10–12 GB VRAM: SDXL Entry Point
12 GB is the practical entry point for SDXL. The RTX 3060 12 GB has become the default recommendation for budget Stable Diffusion builds because it sits right at this threshold.
At 12 GB, you can run SDXL at 1024x1024 with basic workflows. Adding ControlNet or multiple LoRAs pushes you toward the VRAM ceiling. The experience works, but you will hit limits on complex workflows.
For a detailed analysis of what 12 GB cards can and cannot do with SDXL, see our RTX 3060 SDXL real-world limits breakdown.
16–24 GB VRAM: Full SDXL and Training
16 GB (RTX 4070 Ti Super, RTX 4080) gives breathing room for SDXL with ControlNet, multiple LoRAs, and the refiner model. You stop hitting the VRAM wall on most inference workflows.
24 GB (RTX 4090, RTX 3090) is where training becomes practical. LoRA training at 1024x1024 resolution needs 16–24 GB depending on batch size and gradient checkpointing settings. Full fine-tuning requires even more.
The 24 GB VRAM GPU comparison covers the tradeoffs between the RTX 4090 and RTX 3090 at this tier.
Stable Diffusion VRAM Requirements by Model
Different Stable Diffusion versions have different memory footprints. The model architecture determines the baseline VRAM consumption — resolution and batch size scale on top of that.
SD 1.5
SD 1.5 was designed around 512x512 resolution. The UNet model fits within 2–3 GB, and total VRAM usage during inference stays around 4–6 GB at native resolution. Generating at 768x768 pushes usage to 6–8 GB.
SD 1.5 is the lightest model and runs on nearly any modern GPU with 6+ GB VRAM.
SDXL
SDXL uses a larger UNet and operates at 1024x1024 natively. Base inference requires 8–10 GB VRAM. Using the SDXL refiner adds another 2–3 GB. Loading ControlNet or IP-Adapter alongside the base model pushes total consumption to 12–16 GB.
Most SDXL out-of-memory situations occur when users load multiple auxiliary models simultaneously. The base model alone fits in 8 GB with optimization, but the workflow stack determines actual usage. Our SDXL VRAM requirements guide covers this in detail.
SD 3.5
SD 3.5 Large uses a transformer-based architecture that consumes more VRAM than SDXL. Native FP16 inference requires approximately 18 GB for the Large variant. TensorRT with FP8 quantization reduces this to around 11 GB. The Q8 quantized version runs at approximately 16 GB.
SD 3.5 Medium is lighter at around 6 GB, making it accessible on mid-range hardware.
High Resolution and Batch Generation
Generating above native resolution increases VRAM usage non-linearly. Doubling resolution from 1024x1024 to 2048x2048 roughly quadruples memory consumption for the latent space.
Batch generation (producing multiple images simultaneously) multiplies VRAM usage proportionally. A batch of 4 at 512x512 uses similar VRAM to a single image at 1024x1024.
| Scenario | Approximate VRAM |
|---|---|
| SD 1.5 at 512x512, batch 1 | 4–6 GB |
| SD 1.5 at 768x768, batch 1 | 6–8 GB |
| SDXL at 1024x1024, batch 1 | 8–10 GB |
| SDXL + ControlNet + LoRA | 12–16 GB |
| SDXL at 1024x1024, batch 4 | 16–20 GB |
| SD 3.5 Large (FP16), batch 1 | 16–18 GB |
These numbers are approximate ranges based on community-reported usage. Actual consumption varies with specific implementations, optimization settings, and software versions.
Best GPUs for Stable Diffusion
| GPU | VRAM | Best For | Approx. Street Price |
|---|---|---|---|
| RTX 3060 12GB | 12 GB | Budget SDXL inference | $250–300 (used) |
| RTX 4070 Ti Super | 16 GB | SDXL + ControlNet workflows | $750–850 |
| RTX 4090 | 24 GB | All inference + LoRA training | $1,800–2,000 |
| RTX 3090 | 24 GB | Training on a budget (used market) | $700–900 (used) |
| A100 40GB | 40 GB | Professional training, large batches | Cloud only for most users |
| H100 80GB | 80 GB | Full fine-tuning, multi-model | Cloud only |
GPU Capability Matrix
How each GPU handles different Stable Diffusion workloads:
| GPU | SD 1.5 (512x512) | SDXL (1024x1024) | SDXL + ControlNet | LoRA Training |
|---|---|---|---|---|
| RTX 3060 12GB | Comfortable | Tight | Possible with --lowvram | LoRA only (small batch) |
| RTX 4070 Ti Super 16GB | Fast | Comfortable | Yes | LoRA + small DreamBooth |
| RTX 3090 24GB | Fast | Comfortable | Fast | Full LoRA + DreamBooth |
| RTX 4090 24GB | Very fast | Fast | Fast | All training types |
| A100 40GB | Very fast | Very fast | Very fast | Full fine-tuning |
The RTX 3060 12 GB remains the most recommended entry-level card because it is the cheapest GPU that crosses the 12 GB threshold needed for SDXL.
The RTX 4090 dominates both inference speed and training capability at the consumer level. The RTX 3090 vs RTX 4090 comparison covers when the older card still makes sense — primarily when buying used for training workloads where VRAM matters more than speed.
For a broader selection guide, see our best GPU for Stable Diffusion article.
Stable Diffusion Training GPU Requirements
Training consumes more VRAM than inference because the GPU must store model weights, gradients, optimizer states, and activations simultaneously. Inference only holds weights and a single forward pass.
LoRA Training
LoRA (Low-Rank Adaptation) reduces memory requirements by training small adapter layers instead of the full model. For SDXL LoRA training:
- At 512x512: 8–12 GB VRAM is sufficient with gradient checkpointing.
- At 1024x1024: 16–24 GB VRAM. Cards under 24 GB hit VRAM limits quickly as batch size increases.
The RTX 3090 and RTX 4090 (both 24 GB) are the standard choices for LoRA training at home. The RTX 3090 costs roughly half the price used, while the RTX 4090 trains about 2x faster per iteration.
DreamBooth Training
DreamBooth fine-tunes the full UNet with a small dataset. VRAM requirements are higher than LoRA — expect 16–24 GB for SD 1.5 and 24+ GB for SDXL. Gradient checkpointing and mixed precision help, but DreamBooth fundamentally needs more memory than LoRA.
Full Fine-Tuning
Full fine-tuning of the entire model (all weights, no adapter) requires 24–48+ GB depending on the model variant and optimizer. This is impractical on consumer GPUs. Researchers and studios use A100 (40/80 GB) or H100 (80 GB) instances, typically rented from cloud providers.
The economics shift at this point. An RTX 4090 with 24 GB cannot run the workload, regardless of price. A cloud A100 at roughly $0.63/hr or H100 at roughly $2.79/hr becomes the only practical option. Our cloud GPU cost calculator estimates total cost for these workloads.
Running Stable Diffusion on Cloud GPUs
Cloud GPUs solve a specific problem: your local hardware does not have enough VRAM for the workload you need to run.
Common scenarios where cloud makes sense:
- Your card has 8 GB or less and you need SDXL. Renting a 24 GB card is cheaper than buying one for occasional use.
- Training requires 24–48+ GB VRAM that no consumer card provides.
- Batch production at high resolution needs more VRAM than any single consumer GPU offers.
RTX 4090 cloud instances typically cost $0.29–$0.59/hr depending on the provider and pricing model. Vast.ai offers marketplace pricing from $0.29/hr, RunPod Community Cloud charges $0.34/hr, and dedicated providers charge $0.39+/hr for guaranteed availability. See our Vast.ai vs RunPod RTX 4090 price per hour comparison for a provider-by-provider breakdown.
For a broader rental analysis, see our RTX 4090 cloud rental guide. The cloud GPU vs local GPU cost comparison covers when renting overtakes buying.
If you are deciding between investing in hardware or renting, the buy GPU or use cloud decision framework walks through the breakeven calculation.
You can browse available GPU instances and pricing on the SynpixCloud marketplace.
Stable Diffusion vs SDXL GPU Requirements
The jump from SD 1.5 to SDXL is the single biggest VRAM increase in the Stable Diffusion ecosystem. Understanding why helps you plan hardware correctly.
SD 1.5 uses a 860M-parameter UNet operating at 512x512 latent resolution. SDXL uses a 2.6B-parameter UNet at 1024x1024. The larger model requires roughly 2–3x more VRAM for the weights alone. The higher resolution quadruples the latent space memory.
Combined effect: SDXL needs approximately 2–3x the VRAM of SD 1.5 for basic inference. Adding the SDXL refiner, ControlNet, or IP-Adapter pushes the gap to 3–4x.
This is why a GTX 1660 (6 GB) that ran SD 1.5 comfortably cannot handle SDXL at all. The workload crossed a VRAM boundary that no optimization flag can fully compensate for.
Our SDXL VRAM requirements guide covers the memory behavior in detail, including spilling thresholds and optimization strategies for 12 GB cards.
For users running ComfyUI specifically, the ComfyUI memory optimization guide covers node-level strategies to reduce peak VRAM usage across complex workflows.
Conclusion
Stable Diffusion GPU requirements come down to matching VRAM to your workload:
- SD 1.5 inference: 6–8 GB VRAM is sufficient. Most modern GPUs handle this.
- SDXL inference: 12 GB minimum, 16 GB recommended. The RTX 3060 12 GB is the entry point.
- SDXL with complex workflows: 16–24 GB. The RTX 4070 Ti Super or RTX 4090 covers this.
- LoRA training: 24 GB preferred. RTX 3090 (used) or RTX 4090.
- Full fine-tuning: 40–80 GB. Cloud A100 or H100 instances.
The GPU market in 2026 gives you clear tiers. If you know which model and resolution you will run, the VRAM requirement narrows to one or two GPU choices. If your workload exceeds what your local card offers, cloud GPU rental fills the gap without a hardware purchase.
Use the GPU comparison tool to compare specs side-by-side, or the cost calculator to estimate cloud spending for your specific workflow.
Frequently Asked Questions
What GPU is required for Stable Diffusion?
Stable Diffusion requires a GPU with at least 4 GB VRAM for SD 1.5 at 512x512. For SDXL at 1024x1024, 8–12 GB VRAM is the minimum. The RTX 3060 12 GB is the most common entry-level recommendation because it handles both SD 1.5 and basic SDXL workflows.
How much VRAM do I need for Stable Diffusion?
For SD 1.5 inference: 6–8 GB. For SDXL inference: 12–16 GB. For SDXL with ControlNet and LoRAs: 16–24 GB. For LoRA training: 24 GB. For full model fine-tuning: 48 GB or more. Lower VRAM cards can work with optimization flags like --lowvram, but generation speed drops significantly.
Can Stable Diffusion run on a 6 GB GPU?
Yes, but only SD 1.5 at 512x512 resolution. SDXL will not run reliably on 6 GB without heavy optimization that significantly reduces speed and quality. If you need SDXL, 12 GB is the practical minimum.
Is RTX 3060 good for Stable Diffusion?
Yes. The RTX 3060 12 GB is one of the most popular GPUs for Stable Diffusion. It runs SD 1.5 comfortably and handles basic SDXL workflows at 1024x1024. It struggles with complex SDXL pipelines (ControlNet + multiple LoRAs) and is limited for training to small LoRA jobs only.
What GPU do I need for SDXL?
SDXL requires a minimum of 8 GB VRAM with optimization enabled. 12 GB (RTX 3060) is the practical entry point. 16 GB (RTX 4070 Ti Super) is recommended for workflows with ControlNet or the SDXL refiner. 24 GB (RTX 4090 or RTX 3090) is needed for training.
Is cloud GPU worth it for Stable Diffusion?
Cloud GPUs make sense when your local card does not have enough VRAM for SDXL or training workloads. An RTX 4090 cloud instance costs approximately $0.39–$0.80/hr, which is cheaper than buying a $1,800 card if you only need it occasionally. For daily heavy use, local hardware has lower long-term cost.
