lightning-thunder 0.2.6


pip install lightning-thunder

  Latest version

Released: Oct 22, 2025


Meta
Author: Lightning AI
Requires Python: <3.14,>=3.10

Classifiers

Environment
  • Console

Natural Language
  • English

Development Status
  • 3 - Alpha

Intended Audience
  • Developers

Topic
  • Scientific/Engineering :: Artificial Intelligence
  • Scientific/Engineering :: Information Analysis

Operating System
  • OS Independent

Programming Language
  • Python :: 3
  • Python :: 3.10
  • Python :: 3.11
  • Python :: 3.12
  • Python :: 3.13

Give your PyTorch models superpowers ⚡

Thunder Thunder

 

Source-to-source compiler for PyTorch. Understandable. Inspectable. Extensible.

✅ Run PyTorch 40% faster   ✅ Quantization                ✅ Kernel fusion        
✅ Training recipes         ✅ FP4/FP6/FP8 precision       ✅ Distributed TP/PP/DP 
✅ Inference recipes        ✅ Ready for NVIDIA Blackwell  ✅ CUDA Graphs          
✅ LLMs, non LLMs and more  ✅ Custom Triton kernels       ✅ Compose all the above

Thunder is a source-to-source deep learning compiler for PyTorch that focuses on making it simple to optimize models for training and inference.

It provides:

  • a simple, Pythonic IR capturing the entire computation
  • a rich system of transforms that simultaneously operate on the computation IR, the model, and the weights
  • an extensible dispatch mechanism to fusers and optimized kernel libraries

With Thunder you can:

  • profile deep learning programs easily, map individual ops to kernels and inspect programs interactively
  • programmatically replace sequences of operations with optimized ones and see the effect on performance
  • acquire full computation graphs without graph breaks by flexibly extending the interpreter
  • modify programs to fully utilize bleeding edge kernel libraries on specific hardware
  • write models for single GPU and transform them to run distributed
  • quickly iterate on mixed precision and quantization strategies to search for combinations that minimally affect quality
  • bundle all optimizations in composable recipes, so they can be ported across model families

Ultimately, you should think about Thunder as a highly efficient tool to go from “unoptimized” to “optimized”.

If that is of interest for you, read on to Install Thunder and get started quickly.

license CI testing General checks Documentation Status pre-commit.ci status

 

 

Thunder

Quick start

Install Thunder via pip (more options):

pip install lightning-thunder

pip install -U torch torchvision
pip install nvfuser-cu128-torch28 nvidia-cudnn-frontend  # if NVIDIA GPU is present
For older versions of torch

torch==2.7 + CUDA 12.8

pip install lightning-thunder

pip install torch==2.7.0 torchvision==0.22
pip install nvfuser-cu128-torch27 nvidia-cudnn-frontend  # if NVIDIA GPU is present

torch==2.6 + CUDA 12.6

pip install lightning-thunder

pip install torch==2.6.0 torchvision==0.21
pip install nvfuser-cu126-torch26 nvidia-cudnn-frontend  # if NVIDIA GPU is present

torch==2.5 + CUDA 12.4

pip install lightning-thunder

pip install torch==2.5.0 torchvision==0.20
pip install nvfuser-cu124-torch25 nvidia-cudnn-frontend  # if NVIDIA GPU is present
Advanced install options

Install optional executors

# Float8 support (this will compile from source, be patient)
pip install "transformer_engine[pytorch]"

Install Thunder bleeding edge

pip install git+https://github.com/Lightning-AI/lightning-thunder.git@main

Install Thunder for development

git clone https://github.com/Lightning-AI/lightning-thunder.git
cd lightning-thunder
pip install -e .

Hello world

Define a function or a torch module:

import torch.nn as nn

model = nn.Sequential(nn.Linear(2048, 4096), nn.ReLU(), nn.Linear(4096, 64))

Optimize it with Thunder:

import thunder
import torch

thunder_model = thunder.compile(model)

x = torch.randn(64, 2048)

y = thunder_model(x)

torch.testing.assert_close(y, model(x))

Examples

LLM training

Install LitGPT (without updating other dependencies)

pip install --no-deps 'litgpt[all]'

and run

import thunder
import torch
import litgpt

with torch.device("cuda"):
    model = litgpt.GPT.from_name("Llama-3.2-1B").to(torch.bfloat16)

thunder_model = thunder.compile(model)

inp = torch.ones((1, 2048), device="cuda", dtype=torch.int64)

out = thunder_model(inp)
out.sum().backward()

HuggingFace BERT inference

Install Hugging Face Transformers (recommended version is 4.50.2 and above)

pip install -U transformers

and run

import thunder
import torch
import transformers

model_name = "bert-large-uncased"

tokenizer = transformers.AutoTokenizer.from_pretrained(model_name)

with torch.device("cuda"):
    model = transformers.AutoModelForCausalLM.from_pretrained(
        model_name, torch_dtype=torch.bfloat16
    )
    model.requires_grad_(False)
    model.eval()

    inp = tokenizer(["Hello world!"], return_tensors="pt")

thunder_model = thunder.compile(model)

out = thunder_model(**inp)
print(out)

HuggingFace DeepSeek R1 distill inference

Install Hugging Face Transformers (recommended version is 4.50.2 and above)

pip install -U transformers

and run

import torch
import transformers
import thunder

model_name = "deepseek-ai/DeepSeek-R1-Distill-Llama-8B"

tokenizer = transformers.AutoTokenizer.from_pretrained(model_name)

with torch.device("cuda"):
    model = transformers.AutoModelForCausalLM.from_pretrained(
        model_name, torch_dtype=torch.bfloat16
    )
    model.requires_grad_(False)
    model.eval()

    inp = tokenizer(["Hello world! Here's a long story"], return_tensors="pt")

thunder_model = thunder.compile(model)

out = thunder_model.generate(
    **inp, do_sample=False, cache_implementation="static", max_new_tokens=100
)
print(out)

Vision Transformer inference

import thunder
import torch
import torchvision as tv

with torch.device("cuda"):
    model = tv.models.vit_b_16()
    model.requires_grad_(False)
    model.eval()

    inp = torch.randn(128, 3, 224, 224)

out = model(inp)

thunder_model = thunder.compile(model)

out = thunder_model(inp)

Benchmarks

Although is Thunder a tool for optimizing models, rather than an opaque compiler that gets you speedups out of the box, here is a set of benchmarks.

Perf-wise, out of the box Thunder is in the ballpark of torch compile, especially when using CUDAGraphs. Note however that Thunder is not a competitor to torch compile! It can actually use torch compile as one of its fusion executors.

The script examples/quickstart/hf_llm.py demonstrates how to benchmark a model for text generation, forward pass, forward pass with loss, and a full forward + backward computation.

On an H100 with torch=2.8.0 and nvfuser-cu128-torch28 and Transformers 4.55.4 running Llama 3.2 1B we see the following timings:

Transformers with torch.compile and CUDAGraphs (reduce-overhead mode):  521ms
Transformers with torch.compile but no CUDAGraphs (default mode):       814ms
Transformers without torch.compile:                                    1493ms
Thunder with CUDAGraphs:                                                542ms

Plugins

Plugins are a way to apply optimizations to a model, such as parallelism and quantization.

Thunder comes with a few plugins included of the box, but it's easy to write new ones.

  • scale up with distributed strategies with DDP, FSDP, TP ()
  • optimize numerical precision with FP8, MXFP8
  • save memory with quantization
  • reduce latency with CUDAGraphs
  • debugging and profiling

For example, in order to reduce CPU overheads via CUDAGraphs you can add "reduce-overhead" to the plugins= argument of thunder.compile:

thunder_model = thunder.compile(model, plugins="reduce-overhead")

This may or may not make a big difference. The point of Thunder is that you can easily swap optimizations in and out and explore the best combination for your setup.

How it works

Thunder works in three stages:

  1. ⚡️ It acquires your model by interpreting Python bytecode and producing a straight-line Python program

  2. ️⚡️ It transforms the model and computation trace to make it distributed, change precision

  3. ⚡️ It routes parts of the trace for execution

    • fusion (NVFuser, torch.compile)
    • specialized libraries (e.g. cuDNN SDPA, TransformerEngine)
    • custom Triton and CUDA kernels
    • PyTorch eager operations

 

Thunder

 

This is how the trace looks like for a simple MLP:

import thunder
import torch
import torch.nn as nn

model = nn.Sequential(nn.Linear(1024, 2048), nn.ReLU(), nn.Linear(2048, 256))

thunder_model = thunder.compile(model)
y = thunder_model(torch.randn(4, 1024))

print(thunder.last_traces(thunder_model)[-1])

This is the acquired trace, ready to be transformed and executed:

def computation(input, t_0_bias, t_0_weight, t_2_bias, t_2_weight):
# input: "cuda:0 f32[4, 1024]"
# t_0_bias: "cuda:0 f32[2048]"
# t_0_weight: "cuda:0 f32[2048, 1024]"
# t_2_bias: "cuda:0 f32[256]"
# t_2_weight: "cuda:0 f32[256, 2048]"
t3 = ltorch.linear(input, t_0_weight, t_0_bias) # t3: "cuda:0 f32[4, 2048]"
t6 = ltorch.relu(t3, False) # t6: "cuda:0 f32[4, 2048]"
t10 = ltorch.linear(t6, t_2_weight, t_2_bias) # t10: "cuda:0 f32[4, 256]"
return (t10,)

Note how Thunder's intermediate representation is just (a subset of) Python!

Performance

Thunder is fast. Here are the speed-ups obtained on a pre-training task using LitGPT on H100 and B200 hardware, relative to PyTorch eager.

Thunder

Community

Thunder is an open source project, developed in collaboration with the community with significant contributions from NVIDIA.

💬 Get help on Discord 📋 License: Apache 2.0

0.2.7.dev20260201 Feb 01, 2026
0.2.7.dev20260125 Jan 25, 2026
0.2.7.dev20260118 Jan 18, 2026
0.2.7.dev20260111 Jan 11, 2026
0.2.7.dev20260104 Jan 04, 2026
0.2.7.dev20251228 Dec 28, 2025
0.2.7.dev20251221 Dec 21, 2025
0.2.7.dev20251214 Dec 14, 2025
0.2.7.dev20251207 Dec 07, 2025
0.2.7.dev20251130 Nov 30, 2025
0.2.7.dev20251123 Nov 23, 2025
0.2.7.dev20251116 Nov 16, 2025
0.2.7.dev20251109 Nov 09, 2025
0.2.7.dev20251102 Nov 02, 2025
0.2.7.dev20251026 Oct 26, 2025
0.2.6 Oct 22, 2025
0.2.6.dev20251019 Oct 19, 2025
0.2.6.dev20251012 Oct 12, 2025
0.2.6.dev20251005 Oct 05, 2025
0.2.6.dev20250928 Sep 28, 2025
0.2.6.dev20250921 Sep 21, 2025
0.2.6.dev20250914 Sep 14, 2025
0.2.5 Sep 10, 2025
0.2.5.dev20250907 Sep 07, 2025
0.2.5.dev20250831 Aug 31, 2025
0.2.5.dev20250824 Aug 24, 2025
0.2.5.dev20250817 Aug 17, 2025
0.2.5.dev20250810 Aug 10, 2025
0.2.5.dev20250803 Aug 03, 2025
0.2.5.dev20250727 Jul 27, 2025
0.2.5.dev20250720 Jul 20, 2025
0.2.5.dev20250713 Jul 13, 2025
0.2.5.dev20250706 Jul 06, 2025
0.2.5.dev20250629 Jun 29, 2025
0.2.4 Jun 24, 2025
0.2.4.dev20250622 Jun 22, 2025
0.2.4.dev20250615 Jun 15, 2025
0.2.4.dev20250608 Jun 08, 2025
0.2.4.dev20250601 Jun 01, 2025
0.2.4.dev20250525 May 25, 2025
0.2.3 May 23, 2025
0.2.3.dev20250518 May 18, 2025
0.2.3.dev20250511 May 11, 2025
0.2.3.dev20250504 May 04, 2025
0.2.3.dev20250420 Apr 20, 2025
0.2.3.dev20250413 Apr 13, 2025
0.2.3.dev20250406 Apr 06, 2025
0.2.3.dev20250330 Mar 30, 2025
0.2.3.dev20250323 Mar 23, 2025
0.2.2 Mar 20, 2025
0.2.2.dev20250316 Mar 16, 2025
0.2.2.dev20250312 Mar 12, 2025
0.2.2.dev20250309 Mar 09, 2025
0.2.2.dev20250302 Mar 02, 2025
0.2.2.dev20250223 Feb 23, 2025
0.2.2.dev20250216 Feb 16, 2025
0.2.2.dev20250209 Feb 09, 2025
0.2.2.dev0 Mar 20, 2025
0.2.1 Feb 04, 2025
0.2.1.dev20250202 Feb 02, 2025
0.2.0.dev20250126 Jan 26, 2025
0.2.0.dev20250124 Jan 24, 2025
0.2.0.dev20250119 Jan 19, 2025
0.2.0.dev20250112 Jan 12, 2025
0.2.0.dev20250105 Jan 05, 2025
0.2.0.dev20241229 Dec 29, 2024
0.2.0.dev20241222 Dec 22, 2024
0.2.0.dev20241215 Dec 15, 2024
0.2.0.dev20241208 Dec 08, 2024
0.2.0.dev20241201 Dec 01, 2024
0.2.0.dev20241124 Nov 24, 2024
0.2.0.dev20241117 Nov 17, 2024
0.2.0.dev20241110 Nov 10, 2024
0.2.0.dev20241103 Nov 03, 2024
0.2.0.dev20241027 Oct 27, 2024
0.2.0.dev20241020 Oct 20, 2024
0.2.0.dev20241013 Oct 13, 2024
0.2.0.dev20241006 Oct 06, 2024
0.2.0.dev20240929 Sep 29, 2024
0.2.0.dev20240922 Sep 22, 2024
0.2.0.dev20240915 Sep 15, 2024
0.2.0.dev20240908 Sep 08, 2024
0.2.0.dev20240901 Sep 01, 2024
0.2.0.dev20240825 Aug 25, 2024
0.2.0.dev20240818 Aug 18, 2024
0.2.0.dev20240811 Aug 11, 2024
0.2.0.dev20240804 Aug 04, 2024
0.2.0.dev20240728 Jul 28, 2024
0.2.0.dev20240721 Jul 21, 2024
0.2.0.dev20240714 Jul 14, 2024
0.2.0.dev20240707 Jul 07, 2024
0.2.0.dev20240630 Jun 30, 2024
0.2.0.dev20240623 Jun 23, 2024
0.2.0.dev20240616 Jun 16, 2024
0.2.0.dev20240609 Jun 09, 2024
0.2.0.dev20240602 Jun 02, 2024
0.2.0.dev20240526 May 26, 2024
0.2.0.dev20240519 May 19, 2024
0.2.0.dev20240513 May 13, 2024
0.2.0.dev20240505 May 05, 2024
0.2.0.dev20240428 Apr 28, 2024
0.2.0.dev20240421 Apr 21, 2024
0.2.0.dev20240414 Apr 14, 2024
0.2.0.dev20240407 Apr 07, 2024
0.2.0.dev20240404 Apr 04, 2024
0.1.0 Mar 20, 2024
Extras: None
Dependencies:
torch (>=2.7.1)
looseversion (==1.3.0)
lightning-utilities (>0.7.0)
numpy
networkx (>=3.3)
optree (>=0.12.1)
opt_einsum (>=3.3.0)
mpmath (<1.4.0)
dill (>=0.3.8)