LTX Video Model Explained: Lightricks Architecture, Open Weights and Hardware
Most text-to-video systems you can name are rented, not owned. You send a prompt to someone else’s cluster, wait, pay per second of output, and accept whatever usage terms come with it. The LTX video model family from Lightricks is the most prominent attempt to break that pattern: a diffusion transformer that generates video and synchronized audio in one pass, ships its weights publicly, and is engineered to be fast enough that a single high-end GPU is a realistic place to run it.
That matters now because the family has moved quickly. The first model shipped in late 2024 at 2 billion parameters. The current flagship, LTX-2.5, was released in August 2026 as a roughly 22 billion parameter model with native multishot generation, a diffusion-based video decoder and a Gemma 4 12B text encoder. The architecture, the licensing and the hardware arithmetic are all more nuanced than the headlines suggest.
This article separates what Lightricks has documented from what third parties repeat. You will leave with a mechanism-level picture of how the model is built, what “open weights” means under the LTX-2.x Community License, how to reason about VRAM and cost, and where the model is likely to fail.
What this covers: the lineage from LTX Video to LTX-2.5, the dual-stream transformer and decoder design, the training and sampling pipeline, licensing and deployment options, a worked hardware and cost model, failure modes, and a decision matrix against peer models.
Context and Background
Video diffusion has spent three years chasing two conflicting goals: quality and cost. Early systems denoised every frame in pixel space, which made a five-second clip an expensive batch job. The field then moved to latent diffusion, where a variational autoencoder (VAE) compresses video into a small latent grid and a transformer denoises that grid instead. Closed systems such as OpenAI’s Sora pushed this to long, coherent clips; we covered that design in our explanation of the Sora 2 video generation architecture. The open ecosystem lagged on quality and, more painfully, on speed.
Lightricks, the Israeli company behind the Facetune and Videoleap consumer apps, entered with a different bet. Its first model, LTX Video, was described in an arXiv paper (2501.00103) as using a Video-VAE with a 1:192 compression ratio and spatiotemporal downscaling of 32 x 32 x 8 pixels per token. The paper’s headline claim was generating 5 seconds of 24 fps video at 768 x 512 in about 2 seconds on one Nvidia H100, which is faster than real time. The unusual design choice was relocating the patchifying step from the transformer input to the VAE input, so the transformer works in a very compact latent space.
The incumbents in open video are worth naming plainly. Alibaba’s Wan family, Tencent’s HunyuanVideo and Genmo’s Mochi occupy the same open-weights niche, and each makes a different compromise between parameter count, clip length and licence terms. Closed incumbents, including Sora 2 and Google’s Veo line, lead on raw fidelity and duration but cannot be self-hosted. LTX’s differentiator has always been the speed-per-dollar axis rather than the absolute quality ceiling.
There is a second shift that makes the 2026 generation different: audio. Until LTX-2, most open video models were silent, and teams bolted on separate text-to-speech or foley models. LTX-2 generates the soundtrack inside the same denoising process, which changes both the architecture and the evaluation problem. If you work on generative systems more broadly, the same iterative-denoising idea now spans domains as different as text diffusion language models and diffusion policies for robot manipulation, which is why understanding one video model in depth pays off elsewhere.
A note on scope. The writer’s contract for this page is to cover the current shipping version, so the focus is LTX-2.5. Where a fact is documented only for an earlier release (the 2B paper, the LTX-2 paper), I say so. Where Lightricks has not published a number, I say that too, and I flag where independent sources disagree. The review log for this article lists the facts I could not verify from a primary source.
How the LTX Video Model Is Built: Reference Architecture
LTX-2.5 is an asymmetric dual-stream diffusion transformer of roughly 22 billion parameters. One stream denoises video latents, a smaller stream denoises audio latents, and bidirectional cross-attention layers let each stream condition on the other at every block. A Gemma 4 12B language model encodes the prompt, and a diffusion-based decoder turns the final video latents into pixels.

Figure 1: LTX-2.5 generation path. A Gemma 4 12B encoder feeds both transformer streams, which exchange information through cross-attention before separate video and audio decoders.
The figure shows the dataflow, not layer counts. Lightricks has published the stream sizes for LTX-2 (the paper reports a 14 billion parameter video stream and a 5 billion parameter audio stream, 19 billion in total) but the Hugging Face card for the later 2.3 and 2.5 checkpoints states only a 22 billion parameter total. I could not find an official video/audio split for 22B, so treat any specific split you read elsewhere as unverified.
A dual-stream transformer, deliberately asymmetric
Audio and video have very different information density. A second of 24 fps video at delivery resolution carries orders of magnitude more raw data than a second of audio, and the temporal structure differs too: video needs spatial attention across a grid, audio needs long-range sequence modelling. A symmetric design would waste capacity on audio or starve video. The LTX-2 paper’s answer is an asymmetric split, with the larger stream for video and the smaller for audio.
The streams are not independent models glued together. The paper describes bidirectional audio-video cross-attention, temporal positional embeddings, and a cross-modality adaptive layer norm (AdaLN) so that both streams share the same diffusion timestep conditioning. In practice this means a footstep sound and the frame where the foot lands are denoised together, rather than the audio being generated afterwards to match a finished video. That is the mechanism behind lip-sync and foley alignment, and also the reason failures in one modality can leak into the other.
The paper also introduces a modality-aware classifier-free guidance scheme. Classifier-free guidance (CFG) is the standard trick of blending a conditioned prediction with an unconditioned one to strengthen prompt adherence. A modality-aware variant lets the sampler push text adherence and audio-visual alignment with separate strengths instead of a single scalar. I would treat the exact formulation as something to read in the paper rather than infer, since the abstract does not spell it out.
The text encoder and why it changed
LTX-2 used a multilingual text encoder. LTX-2.5 is documented as using a custom Gemma 4 12B text encoder, paired with a learned projection and an optional prompt enhancer. The practical reason to move to a larger language-model backbone is prompt fidelity over long, structured prompts, which is exactly what multishot generation needs: a prompt that describes three shots, a character and a line of dialogue has to survive encoding intact.
This choice has a cost that rarely makes the marketing copy. A 12B text encoder is itself a large model. In bf16 that is roughly 24 GB of weights before the transformer is loaded. Pipelines therefore load the encoder, compute the embeddings, and release or offload it before denoising begins. That sequencing is one reason the headline “16 GB VRAM” figure depends on offload flags, which I quantify later.
The video latent space and the decoder
The original LTX Video paper’s 1:192 compression is the clearest published description of the family’s philosophy: make the latent grid so small that the transformer’s quadratic attention cost stays manageable, then pay for fidelity in the decoder. The 2024 paper had the VAE decoder do the final denoising step in pixel space, avoiding a separate upsampler.
LTX-2.5 pushes that idea further with what Lightricks calls a diffusion video decoder. Instead of a deterministic VAE reconstruction, the decoder itself is generative, which restores high-frequency detail such as faces and small text that heavily compressed latents lose. The ltx.io page positions it as improving motion quality, and the independent summaries describe sharper faces and text. I did not find a published ablation quantifying the improvement, so the claim is qualitative.
The Hugging Face card lists two video VAE options in the inference stack: a DiffVAE for higher quality and a Conv VAE for speed, plus a spatial upscaler and a temporal upscaler, each by a factor of two. An audio VAE with an integrated vocoder handles the sound path. The card does not give the 2.5 latent compression ratios, so the 32 x 32 x 8 figure should be read as the 2024 model’s number, not a verified 2.5 specification.
Two pipelines, one checkpoint family
For deployment, the ltx-pipelines package exposes several entry points. The DistilledPipeline is the fastest text-to-video and image-to-video path. The DFRPipeline, built around Diffusion Fidelity Rendering, is the production-quality variant that, per the repository, needs longer runtime and more VRAM. Other pipelines cover two-stage text-to-video, keyframe interpolation, audio-to-video, retakes, HDR in-context LoRA and dubbing.
Diffusion Fidelity Rendering is described by Lightricks as allocating rendering compute according to scene complexity. I read that as adaptive spending: simple, static regions need fewer refinement operations than a crowded, fast-moving scene. The exact algorithm is not documented in the sources I could reach, so I will not speculate on whether it is adaptive step counts, adaptive resolution or something else.
Lineage: From 2 Billion Parameters to LTX-2.5

Figure 2: LTX release lineage. Dates follow the Wikipedia version history and the Comfy blog; parameter counts are those published by Lightricks.
The Wikipedia version history gives the cleanest timeline, cross-checked against vendor material where possible. LTX Video arrived in November 2024 at 2 billion parameters. A 13 billion parameter model, LTXV-13b, followed in May 2025, with Lightricks claiming large speed gains. A July 2025 update was reported as breaking the 60-second generation barrier. LTX-2 was announced in October 2025 with unified audio-video generation, and its weights and tooling were released openly in January 2026, alongside the arXiv paper (2601.03233).
LTX-2.3 followed in March 2026 and is now described by Lightricks as the previous generation, still supported, with open weights, synchronized audio and native portrait video. LTX-2.5 launched in August 2026. Sources disagree by a day on the date: the ComfyUI Wiki news item is dated August 11, 2026, while the official Comfy blog post on day-zero support gives August 12, 2026. Lightricks’ own model pages did not state a launch date in the content I could retrieve, so I would say “mid-August 2026” and cite both.
What changed in 2.5 relative to 2.3, by the vendor’s own list: native multishot generation that holds character, environment, lighting and voice across connected shots; an optional duration predictor that chooses the clip length from the requested action; the diffusion video decoder; Diffusion Fidelity Rendering; and higher-end HDR workflows, including native 4K HDR, SDR-to-HDR conversion and 16-bit EXR input and output. The Hugging Face card notes that LoRAs trained on LTX-2.3 mostly run on 2.5 unchanged, though it advises validating before production.
For readers following the series, this model class sits alongside the closed systems in our Sora 2 piece, and it shares a conceptual thread with the world-model work covered in world action models versus vision-language-action systems: learning a generative model of visual dynamics conditioned on language.
Deeper Analysis: Sampling, Memory Arithmetic and Cost
The sampling pipeline in practice

Figure 3: Generation pipeline. The distilled path trades steps for speed; the DFR path spends more compute. Both converge on the diffusion decoder, with optional upscaling stages. This is a pipeline view, not a benchmark chart.
Because Lightricks has not published a benchmark table for 2.5 that I could verify, this figure shows the pipeline rather than a bar chart of scores. The structure is common to modern open video systems. A first pass generates latents at reduced resolution, a spatial upscaler doubles them, a refinement pass repairs the detail the upscaler invented, and the decoder converts latents to pixels. A temporal upscaler can double frame rate afterwards.
The distilled checkpoint is the key to the speed claims. The model card describes a fixed 8-step schedule with CFG equal to 1. Standard diffusion with classifier-free guidance runs the network twice per step, once conditioned and once unconditioned, so setting CFG to 1 removes that second pass. Eight steps at one forward pass each is roughly a 6x reduction in network evaluations relative to a typical 25-step CFG schedule. That arithmetic is illustrative, since the dev checkpoint supports variable steps and Lightricks does not prescribe 25, but it shows why distillation compounds: fewer steps and no guidance doubling.
The cost of distillation is flexibility and sometimes diversity. A few-step student model learns to imitate a many-step teacher, and it generally has less room for negative prompts, guidance tuning and fine control. The dev checkpoint remains the one to fine-tune and the one that supports variable steps, which is why the repository keeps both.
Why the token count drives everything
Transformer cost in video diffusion is dominated by the number of latent tokens, because self-attention scales with the square of the sequence length. The following is an illustrative calculation using the 2024 paper’s compression (32 x 32 x 8 pixels per token) as a stand-in. LTX-2.5’s exact latent geometry is not documented in the sources I could reach, so these numbers show the scaling behaviour, not 2.5’s real token counts.
| Clip | Frames | Latent grid (w x h x t) | Tokens | Attention cost vs 768×512 |
|---|---|---|---|---|
| 768 x 512, about 5 s at 24 fps | 121 | 24 x 16 x 16 | 6,144 | 1x |
| 1920 x 1088, about 5 s at 24 fps | 121 | 60 x 34 x 16 | 32,640 | about 28x |
| 3840 x 2176, about 5 s at 24 fps | 121 | 120 x 68 x 16 | 130,560 | about 452x |
The frame count follows the documented constraint that frames must satisfy N mod 8 equals 1, so 121 frames give (121 minus 1) divided by 8 plus 1, or 16 latent time steps. The pixel dimensions are both divisible by 32, matching the documented requirement. Moving from 768 x 512 to 4K multiplies the token count by about 21 and the quadratic attention term by roughly 450. Real implementations use memory-efficient attention kernels and the model has non-attention costs that scale linearly, so wall-clock time grows by less than the raw quadratic figure, but the direction is unmistakable.
This is why every serious pipeline for 4K uses a staged approach: generate at a lower resolution, upscale in latent space, refine. It is also why the hosted API charges much more for 4K ($0.30 per second on the Fast tier against $0.09 at 720p) and why claims of faster-than-real-time generation always specify a resolution and hardware. The headline in Lightricks material is a 10-second clip at 720p in 6.8 seconds on two NVIDIA GB200 GPUs. That is a vendor claim on very expensive datacenter hardware, and it does not transfer to a consumer card.
VRAM: what 16 GB really means
The advertised minimum is 16 GB of VRAM, which appears on the ltx.io page and in secondary coverage. The Hugging Face model card itself does not state a numeric VRAM requirement, and the repository’s guidance is qualitative: use the fp8-cast quantization with CPU or disk offload for low-memory setups. So the 16 GB number is best read as “can run with aggressive offload”, not “runs comfortably”.
Weight arithmetic makes the picture concrete. These are illustrative calculations from parameter counts, not measured usage.
| Component | Parameters | bf16 (2 bytes) | fp8 or int8 (1 byte) | 4-bit class (about 0.5 byte) |
|---|---|---|---|---|
| Diffusion transformer | about 22B | about 44 GB | about 22 GB | about 11 GB plus scales |
| Gemma 4 text encoder | 12B | about 24 GB | about 12 GB | about 6 GB plus scales |
A bf16 transformer alone is around 44 GB, which exceeds every consumer GPU and fits only on workstation or datacenter cards. At one byte per parameter it is around 22 GB, still above 16 GB. Reaching a 16 GB card therefore needs weights streamed in from system RAM or disk during the forward pass, which is what the offload flags do. The NVFP4 and int8 convrot checkpoints published for ComfyUI shrink the footprint further, though the Hugging Face card says NVFP4 needs the ltx-kernels package or Blackwell-class hardware, and the int8 variant is ComfyUI-only and incompatible with the plain PyTorch path.
Activations add to weights. At the token counts above, the attention and feed-forward activations for a 4K clip can be significant even with efficient kernels, which is why the repository warns that the DFR pipeline needs more VRAM than the distilled one. Offload trades VRAM for time: each block’s weights cross the PCIe bus, and for a 22B model that traffic can dominate runtime. A card with 24 GB or more, or a system with fast unified memory, will behave very differently from a 16 GB card at the same nominal support level. I have not measured this; it follows from the arithmetic and from the repository’s own flags.
Training memory is a separate budget
The vendor material cites different figures for fine-tuning: about 80 GB for standard training and about 32 GB for a low-memory configuration, via the ltx-trainer package that supports LoRA, full fine-tuning and in-context LoRA (IC-LoRA). These come from a third-party review summarising Lightricks documentation, so check the trainer README before provisioning hardware. The qualitative point stands: inference on a consumer card is feasible, full fine-tuning is not, and LoRA is the practical path for most teams.
API economics versus owning a GPU
The hosted pricing published on ltx.io is per second of generated video. For the Fast endpoint: $0.09 at 720p, $0.13 at 1080p, $0.19 at 1440p and $0.30 at 4K. For Pro: $0.12 at 720p and $0.17 at 1080p. A 10-second 1080p Fast clip is therefore $1.30, and a 4K clip of the same length is $3.00. Prices change, so confirm at docs.ltx.io/pricing.
An illustrative break-even sketch, with the GPU rate left as a variable because I have no verified figure: if a rented or owned GPU costs G dollars per hour all-in and your pipeline produces C finished clips per hour at the target resolution, local cost per clip is G divided by C. At 1080p Fast, an API clip of 10 seconds costs $1.30. If you iterate heavily, discarding most generations, the effective cost per accepted clip multiplies by your rejection rate on either route, so the break-even depends more on hit rate and utilisation than on list price. Owning a GPU wins when it is busy; the API wins when demand is spiky, and it also removes the VRAM and offload engineering.
Licence, Access and Deployment

Figure 4: Deployment decision flow for the LTX video model, from the revenue threshold through hardware to checkpoint choice.
What “open weights” means here
The weights are downloadable from Hugging Face and the code is on GitHub, but the licence is not a standard open-source licence. The LTX-2.x Community License Agreement allows free commercial and production use for organisations under 10 million US dollars in annual revenue, and entities at or above that threshold need a paid commercial agreement. The Hugging Face card also warns that fine-tune transfers may require paid licensing, which is the clause legal teams should read first. The binding terms are in the repository’s LICENSE files, and I have not reproduced them, so read them directly before shipping a product.
This matters for the vocabulary. Under the Open Source Initiative’s definition, a revenue cap is a restriction on field of use, so the family is better described as open weights with a community licence than as open source. Marketing pages on ltx.io describe the licence as permissive with no mandatory branding, which is accurate about watermarks and attribution, but not a statement that the 10 million threshold disappears. Earlier releases carried different terms, so do not assume the LTX Video 2B licence applies to the current model.
Checkpoint choice
The model card lists a distilled bf16 transformer (fixed 8-step schedule, CFG of 1), a dev bf16 transformer that is trainable and supports variable steps, an int8 convrot variant for ComfyUI, and an NVFP4 variant. A distilled LoRA and a duration-head patch are also listed. Choose the distilled checkpoint for throughput, iteration and interactive tooling. Choose dev for fine-tuning, for guidance control and when you need the maximum quality the pipeline can deliver.
Inference stacks
Three routes exist. The ltx-pipelines Python package is the reference implementation; it is installed with uv and has an optional natten extra tuned for Linux and CUDA. ComfyUI has had day-zero support since version 0.32.0 with templates for text-to-video, image-to-video and first-frame-to-last-frame workflows. Diffusers support exists through a separate package named LTX-2.5-Diffusers, according to the model card. The 2.3 card lists Python 3.12 or newer, CUDA above 12.7 and PyTorch around 2.7 as the tested environment, and mentions Apple device mapping through MPS; I could not confirm the same matrix for 2.5, so pin versions from the 2.5 README.
Resolution, frame rate and length
Documented constraints are consistent across cards: width and height divisible by 32, and a frame count that equals a multiple of 8 plus 1. The example shown on the card runs at 24 fps. Official pages list output frame rates of 24, 25, 48 and 50 fps, and resolutions from 720p up to native 4K, with 4K and 1440p generation on the Fast tier. Duration depends on tier and endpoint: up to 20 seconds on the Fast tier and up to 10 seconds on Pro per the Comfy blog, while a third-party review describes 6 to 20 seconds depending on model, resolution, frame rate and endpoint. For the open weights run locally, the practical limit is your VRAM and patience rather than a hard cap, and Lightricks does not document a maximum frame count for self-hosted use.
The earlier claim of generating more than 60 seconds dates from July 2025 and applied to the older model; I would not extrapolate it to LTX-2.5, whose official materials emphasise clips of 10 to 20 seconds and multishot sequences instead.
Decision matrix against peers
The matrix below is qualitative and reflects documented properties rather than my own benchmark runs. Peer details are from general knowledge of the models and should be re-verified at the time you choose.
| Use case | LTX-2.5 | Wan-class open models | Closed APIs such as Sora 2 or Veo |
|---|---|---|---|
| Local iteration on one GPU | Strong: distilled 8-step path, quantized checkpoints | Viable, typically slower per clip | Not possible |
| Video with synchronized audio in one pass | Native, single model | Often needs a separate audio step | Native in newer closed models |
| Fine-tuning for a brand or character | LoRA and IC-LoRA via ltx-trainer | Community LoRA ecosystems | Limited or unavailable |
| Longest, most coherent hero shots | Moderate: 10 to 20 second clips, multishot | Moderate | Generally stronger |
| Licence flexibility at scale | Revenue threshold applies | Varies by model | API terms only |
| Data residency and offline use | Full control | Full control | None |
The thesis I would defend, as opinion: the LTX video model is best understood as an infrastructure choice rather than a quality contest. If your constraint is cost per iteration, control over weights or offline operation, it is the leading open candidate. If your constraint is the single most convincing 30-second shot, rented frontier systems still deserve the first look.
Trade-offs, Gotchas, and What Goes Wrong
Documented limitations. The Hugging Face cards for 2.3 and 2.5 are blunt: the model is not intended or able to provide factual information, it may amplify societal biases, prompt following depends heavily on prompting style, adherence is imperfect, and it can generate inappropriate or offensive content. The 2.3 card adds that audio quality is lower when generating without speech. Treat those as the vendor’s own admissions, and plan moderation accordingly if you expose the model to users.
Joint generation couples failure modes. Because audio and video denoise together, a prompt that confuses one stream can degrade the other. Early reviews of the LTX-2 generation noted occasional lip-sync and motion-tracking inconsistencies in long scenes and difficulty with precise text rendering inside video. The diffusion decoder in 2.5 targets the text and face problem, but I found no independent test showing it is solved, so verify with your own prompts, particularly signage and captions.
Multishot consistency is a claim, not a guarantee. Native multishot generation promises that character, environment, lighting and voice persist across cuts. Identity drift across shots is the classic failure of every video model, and the sources I could reach describe the feature without publishing a quantitative consistency metric. Test it on your characters with your prompts before you build a workflow around it.
The 16 GB headline hides latency. Running a 22B transformer and a 12B text encoder on a 16 GB card requires streaming weights, and the speed claims were measured on datacenter hardware. If your plan assumes interactive generation on a gaming card, benchmark one clip end to end first. Also remember that quantized formats are tied to stacks: NVFP4 depends on Blackwell-class GPUs or the ltx-kernels package, and the int8 convrot checkpoint works only in ComfyUI.
Fine-tune licensing can surprise teams. The card’s note that fine-tune transfers may need paid licensing means a LoRA you train and distribute commercially may inherit constraints. Revenue thresholds are also evaluated at the entity level, which matters for subsidiaries and agencies serving large clients. Get a legal read rather than relying on a summary, including this one.
Version churn. The family moved from 2.3 to 2.5 in about five months. LoRAs mostly carry over, but pipelines, checkpoints and ComfyUI versions change. Pin the repository commit, the checkpoint hash and the ComfyUI version, and re-run a fixed regression prompt set whenever any of them changes.
Evaluation is thin. I could not locate a vendor-published benchmark table for 2.5 against named peers on a recognised video benchmark. The LTX-2 paper claims state-of-the-art quality among open-source systems and comparability with proprietary models at lower cost, but that is the authors’ claim. Independent leaderboards for video change often, and human-preference scores are noisy. Build a small internal evaluation set that reflects your content, and score it blind.
Provenance and safety. Generative video raises authenticity questions regardless of the licence. If you ship generated clips, consider embedding content credentials and disclosure, and review the acceptable-use terms in the licence. The model card does not describe a watermark scheme that I could verify, and the vendor says there is no mandatory branding.
Practical Recommendations
Start by deciding whether the licence fits. If your organisation is at or above 10 million dollars in annual revenue, or you plan to distribute fine-tunes commercially, talk to Lightricks before you build. Under the threshold, you can use the weights for commercial work at no cost, subject to the licence text.
Then match the route to your workload. For occasional clips or spiky demand, the hosted Fast tier is the cheapest way to learn what the model does, and its per-second pricing makes budgeting simple. For sustained iteration, a local deployment on a 24 GB or larger GPU with the distilled checkpoint is the sensible default, using fp8-cast quantization and offload only when you must. Reserve the dev checkpoint and the DFR pipeline for final renders and fine-tuning.
Keep prompts structured. The Gemma-based encoder and the prompt enhancer are designed for long, descriptive prompts, and the card stresses that prompting style strongly influences adherence. Write shots as separate sentences with subject, action, camera and lighting, and add dialogue explicitly if you want speech.
- Read the LTX-2.x Community License and confirm the revenue threshold and fine-tune clauses with legal counsel.
- Pick a route: hosted Fast or Pro, or local distilled or dev checkpoint.
- Benchmark one 10-second clip end to end on your hardware at your target resolution before committing.
- Pin ltx-pipelines, checkpoint hashes and ComfyUI (0.32.0 or later) versions.
- Build a fixed regression prompt set covering text on screen, faces, multishot identity and audio sync.
- Add moderation and disclosure steps before exposing generation to end users.
- Track the pricing page and the model card for changes, since both have moved between releases.
Frequently Asked Questions
What is the LTX video model?
LTX is a family of text-to-video diffusion transformer models from Lightricks. The first model appeared in November 2024 at 2 billion parameters. LTX-2 added joint audio and video generation, and the current LTX-2.5, released in August 2026, has about 22 billion parameters. It supports text-to-video, image-to-video, video-to-video and audio-to-video, produces synchronized sound in the same pass, and ships open weights under a community licence alongside a paid API.
Is the LTX video model open source?
It is open weights, not open source in the strict sense. The weights are public on Hugging Face and the code is on GitHub, but the LTX-2.x Community License allows free commercial use only for organisations under 10 million dollars in annual revenue. Larger entities need a paid agreement, and fine-tune transfers may require paid licensing. Read the LICENSE files in the repository for the binding terms before shipping a commercial product.
How much VRAM do you need to run LTX-2.5 locally?
Lightricks advertises a minimum of 16 GB of VRAM, but the model card gives no numeric requirement and the repository recommends fp8-cast quantization with CPU or disk offload for low-memory setups. By arithmetic, a 22 billion parameter bf16 transformer is about 44 GB of weights, so 16 GB cards depend on quantization and offload and will be slower. A 24 GB or larger card is more comfortable. These are estimates, not measurements.
How long and how sharp can LTX video clips be?
Official pages list resolutions from 720p up to native 4K at 24, 25, 48 or 50 fps. Duration depends on the tier: the Comfy blog reports 2 to 20 seconds on the Fast variant and 2 to 10 seconds on Pro. Frame counts must equal a multiple of 8 plus 1, and width and height must be divisible by 32. For self-hosted weights, no hard maximum is documented, so memory and time are the practical limits.
Does LTX generate audio as well as video?
Yes. Since LTX-2 the model generates synchronized speech, ambience and foley in the same denoising process. The architecture uses an asymmetric dual-stream transformer, with a larger video stream and a smaller audio stream joined by bidirectional cross-attention. Audio is decoded by an audio VAE with an integrated vocoder. The model card notes that audio quality is lower when generating without speech, so test non-speech soundscapes carefully.
How does LTX-2.5 differ from LTX-2.3?
According to Lightricks, LTX-2.5 adds native multishot generation that keeps characters and settings consistent across cuts, an optional duration predictor, a diffusion video decoder for sharper detail, Diffusion Fidelity Rendering that allocates compute by scene complexity, and stronger HDR workflows. It also uses a Gemma 4 12B text encoder. LoRAs trained on 2.3 mostly work unchanged on 2.5, though the card advises validating before production use.
Further Reading
- OpenAI Sora 2 explained: video generation architecture, the closed-model counterpart to this open family.
- Diffusion LLMs and text diffusion architecture, the same iterative-denoising idea applied to language.
- Diffusion policy for robot manipulation and imitation learning, diffusion used to generate actions rather than pixels.
- World action models versus VLA systems, generative models of dynamics in robotics.
- LTX-2.5 model card on Hugging Face, the primary technical reference.
- Lightricks LTX-2 repository on GitHub, pipelines, trainer and installation notes.
References
- HaCohen et al., “LTX-Video: Realtime Video Latent Diffusion,” arXiv:2501.00103. https://arxiv.org/abs/2501.00103
- “LTX-2: Efficient Joint Audio-Visual Foundation Model,” arXiv:2601.03233. https://arxiv.org/abs/2601.03233
- Lightricks, LTX-2.5 model card, Hugging Face. https://huggingface.co/Lightricks/LTX-2.5
- Lightricks, LTX-2.3 model card, Hugging Face. https://huggingface.co/Lightricks/LTX-2.3
- Lightricks, LTX-2 repository (ltx-core, ltx-pipelines, ltx-trainer). https://github.com/Lightricks/LTX-2
- LTX model pages and pricing, ltx.io. https://ltx.io/model/ltx-2-5
- ComfyUI blog, “LTX-2.5 Day-0 Support in ComfyUI.” https://blog.comfy.org/p/ltx-25-day-0-support-in-comfyui
- ComfyUI Wiki, “LTX-2.5: Lightricks Launches 22B Open Video Model for ComfyUI,” August 11, 2026. https://comfyui-wiki.com/en/news/2026-08-11-ltx-2-5-open-weights-release
- Wikipedia, “LTX (text-to-video model),” version history. https://en.wikipedia.org/wiki/LTX_(text-to-video_model)
By Riju – about
