FLUX 3 Action: Black Forest Labs Enters Robot Control with a World Action Model

FLUX 3 Action: Black Forest Labs Enters Robot Control with a World Action Model

FLUX 3 Action: Black Forest Labs Enters Robot Control with a World Action Model

The company that made its name with image generators has shipped a robot policy. FLUX 3 Action is a 7-billion-parameter model from Black Forest Labs (BFL) that looks at camera frames, joint positions and a text instruction, then predicts the next 2.13 seconds of motor commands and the video of what should happen while those commands run. BFL published the weights on Hugging Face on September 24, 2026, and reports a leaderboard score of 42.92% on RoboLab-120, ahead of a 16B model from NVIDIA and a 3.3B model from Physical Intelligence.

That matters because it moves the “video model as robot brain” idea from research demos to a checkpoint you can download, fine-tune on 24 GB of VRAM, and benchmark yourself. It also arrives with caveats that most launch coverage skips: a custom licence, no built-in safety limits, and a benchmark that is simulation-only.

This post explains how the model is built, what the numbers do and do not show, how it differs from vision-language-action (VLA) models, and how to fit it into a robot software stack.

What this covers: the architecture, the training mixture, the RoboLab-120 methodology, latency and hardware maths, licence terms, failure modes, and a decision matrix against VLAs and other world action models.

Context and Background: From Image Generator to Robot Policy

For about two years the dominant recipe for generalist robot policies was the vision-language-action model. A VLA starts from a pretrained vision-language model, adds an action head, and is trained on robot demonstrations to output actions given images and an instruction. Physical Intelligence’s pi0 and pi0.5, NVIDIA’s GR00T family and Google’s Gemini Robotics all follow variations of that pattern. Our comparison of VLA models covers the lineage in detail.

A second family has grown alongside it: the world action model (WAM). Instead of borrowing a language model’s understanding of the world, a WAM starts from a video generator, which has already learned how scenes evolve over time, and teaches it to emit actions. The bet is that predicting pixels forward forces the network to internalise contact, occlusion, gravity and object permanence, which language pretraining teaches only indirectly. NVIDIA’s Cosmos line, Cosmos Policy variants and several academic systems sit in this camp. The trade-offs are explored in world action models vs VLA: Cosmos 3, VLA-JEPA and FastWAM, and the underlying “world model” idea in our Cosmos architecture reference.

BFL is a newcomer to robotics but not to video and image synthesis. Its FLUX image models (see our FLUX architecture and benchmark review) are built on rectified-flow transformers. FLUX 3 is described by BFL as a multimodal backbone trained on image, video and audio data. FLUX 3 Action reuses that backbone and adds robot state and action pathways. Reporting from MarkTechPost and The Decoder agrees with BFL’s model page on the core facts, though the secondary outlets differ on some latency and licence details, which this post flags where relevant. Release date: September 24, 2026.

Why does this matter for industrial and digital-twin work? Because a policy that also predicts video is, in effect, a learned simulator running inside the controller. That is a natural fit for the industrial world-model thinking that digital-twin teams already use, though the predicted video is a hypothesis, not ground truth.

Reference Architecture: How FLUX 3 Action Turns Pixels and State into Motion

Direct answer: FLUX 3 Action is a 7B world action model. It encodes text, camera frames and robot joint state into one token sequence, runs it through FLUX 3 backbone layers, and decodes two outputs at once: future video tokens that become predicted frames, and action tokens that become a 32-step, 15 Hz chunk of robot commands. Both are generated with flow matching.

FLUX 3 Action world action model architecture showing text, camera and state tokens feeding a joint backbone that decodes video and action

Figure 1: FLUX 3 Action data path. Three input streams become one token sequence; the backbone decodes predicted video and an action chunk in the same pass. The safety layer is not part of the model, which is a deliberate omission by BFL.

Figure 1 summarises the path. Text instructions are encoded by Qwen3-VL-4B, which the DROID model card lists under Apache-2.0. Camera frames and joint state are tokenised. BFL’s description is that the backbone layers “carry task, history, state, future, and action features”. Future tokens decode to predicted video frames, and action tokens decode to robot actions. In practice you can consume only the action output and ignore the predicted frames.

One backbone, two heads of output

The central design choice is joint prediction. In a typical VLA the action head is a small module attached to a language model’s final layers. In FLUX 3 Action the video and action streams share the backbone and are denoised together. Action tokens can attend to future-video tokens during generation, so the plan for the arm is conditioned on a simultaneous imagined outcome.

BFL trains with flow matching rather than classical diffusion. In flow matching, the network learns a velocity field that transports noise to data along a time-indexed path, and sampling integrates that field in a small number of steps. This is the same family as the FLUX image models, which is why BFL can reuse infrastructure and why step and guidance distillation, familiar from image generation, port over cleanly.

BFL reports some specific training-recipe details: a logit-logistic timestep distribution with scale 0.75, a shift of 42, an action scaling factor of 2, and an online power EMA with relative sigma 0.10. Those are tuning facts rather than architecture, but they signal that the action stream needed different noise scheduling from images. High shift values push more training weight toward high-noise timesteps, where the coarse structure of a trajectory is decided.

The action space: EE50, joints, and gaming

The model must serve many embodiments, so BFL defines shared action spaces. The main one, called EE50, is a 50-dimensional end-effector representation. Per arm it uses 25 dimensions: 3 for translation, 6 for rotation in the continuous 6D format from Zhou et al. (2019), and 16 for hand-pose angles or gripper state. Two arms give 50. A joint-space variant uses 14 dimensions, 7 per arm (six joint values plus a gripper). A gaming space has 64 dimensions: 2 mouse axes, 2 mouse buttons and 60 keyboard buttons.

The gaming space looks odd in a robotics model but is central to the data story below. Action rates in the datasets range from 5 Hz to 30 Hz, and the model normalises to chunks of 32 actions. At about 15 Hz, one chunk covers 32 / 15 = 2.13 seconds.

The released DROID checkpoint is more concrete. Its inputs are three RGB views (wrist, left and right, at 360 by 640), seven arm joint positions plus a gripper fraction, and text. Its outputs are absolute joint commands followed by a gripper closed fraction (0 open, 1 closed), 32 actions at 15 Hz. BFL states there are no built-in velocity, force or workspace limits. That sentence is the most important line in the model card for anyone connecting the policy to hardware.

Why an action chunk of 2.13 seconds

Chunked prediction is standard in modern policies, including the diffusion policy and pi0 families, because it amortises inference cost and smooths trajectories. A 2.13-second horizon is long enough to cover a reach-and-grasp primitive but short enough that a fresh observation can correct drift. Section “Latency, Hardware and Cost” works through how chunking lets a slow policy drive a fast controller.

The model ships in three inference recipes: a base model (4 steps, guidance off), a guidance-distilled variant (4 steps with guidance baked in), and a step-distilled variant (1 step). Each is offered in BF16 and in FP8 with rowwise dynamic scaling (“FP8r”). LeRobot integration is provided for the BF16 packages; the FP8r variants need BFL’s own flux-action package, which is Apache-2.0 licensed.

Deeper Analysis: Training Data, RoboLab-120, and What the Scores Mean

Training: video first, then actions from five sources

BFL says pretraining used image, video and audio data, with video making up more than 95% of tokens. That inheritance from FLUX 3 is the world-model half of the story. The action half arrives in a midtraining stage that mixes five sources. BFL reports both the sample share and the token share of each:

Midtraining source Share of samples Share of tokens
Joint video/audio (retained pretraining data) 36.95% 39.00%
Gaming recordings with keyboard and mouse 19.55% 21.85%
Egocentric human hand video 13.54% 13.78%
Handheld gripper data 14.03% 13.99%
Teleoperation across 14 embodiments 15.93% 11.37%

Source: BFL model page. The five rows sum to 100% of samples.

Two things stand out. First, only about 16% of samples come from conventional robot teleoperation. The rest is video with some kind of action label, or unlabelled video kept to preserve generative quality. Second, the retained 36.95% of pretraining data is a regulariser: it stops the backbone forgetting how to model video while it learns actions. MarkTechPost describes the split as roughly 37% pretraining data and 63% action-aligned video.

Gaming data teaches a general skill: mapping a control signal to a change in what the camera sees. A keyboard press that moves an avatar is structurally similar to a joint command that moves an arm. Egocentric hand video and handheld grippers (a person holds a gripper-like tool with a camera on it) supply manipulation diversity without the cost of robot time. The bet is that the shared EE50 space lets those sources transfer to robot arms.

After midtraining, the model is fine-tuned per embodiment. For DROID, BFL reports 20k fine-tuning steps. The recipe uses 1k frozen steps and 2k linear warmup on the midtraining weights, 1k warmup for action heads, and joint-space fine-tuning with a batch size of 2,048, which BFL says outperformed fine-tuning in EE50 space. The secondary outlet DataNorth reports that adapting to a new arm takes roughly 200 recorded demonstrations. BFL’s page does not state that figure, so treat it as a secondary claim.

The pretraining ablation

The most useful number in the launch may be an ablation. In BFL’s account, without the FLUX 3 pretraining and action midtraining, training on DROID alone left RoboLab performance below 1%. With pretraining, the same protocol reached 11.6%, and kept improving slowly to 12.4%. Note what this is: a controlled comparison under a fixed fine-tuning budget, and the absolute numbers (11.6% and 12.4%) are far below the 42% headline. They come from a different, cheaper protocol. The message is about relative value: the world-model pretraining is what lifts a policy from unusable to functional. It is a claim from the vendor’s own report, and independent replication is still to come.

The RoboLab-120 Benchmark: What It Measures

RoboLab is an NVIDIA-led benchmark described in an arXiv paper (2604.09860) by researchers including Xuning Yang, Rishit Dagli, Alex Zook and Jonathan Tremblay, with collaborators at the University of Toronto and University of Sydney. It is built on Isaac Lab and uses a Franka Panda arm with a Robotiq gripper, matching the DROID hardware configuration. RoboLab-120 is its curated set of 120 tabletop manipulation tasks.

Task facts from the paper and project page: 65 simple, 38 moderate and 18 complex tasks; an average of 2.02 subtasks and 9.0 objects per task; a mean difficulty score of 2.90; and only 68.7% of benchmark objects appear in the DROID training vocabulary, so policies face some novelty. Tasks are organised along three competency axes: visual (colour, semantics, size), relational (counting, spatial and conjunction reasoning) and procedural (affordances, reorientation, stacking). Perturbation axes include camera pose, object placement, lighting, background texture and instruction phrasing.

The design intent is to test policies trained on real-world data without simulation co-training. Metrics beyond success rate include graded scores, motion smoothness (SPARC), path length and event tracking such as collisions. The authors report that procedural reasoning is the hardest axis, that vague instructions degrade performance, and that policies depend heavily on wrist-camera calibration.

The scoreboard, and why two numbers exist

The launch materials contain two FLUX 3 Action figures. Table 1 on BFL’s page lists 42.92% for the leaderboard entry. The running text and multi-seed evaluation give 42.2% (plus or minus 0.36) for the guidance-distilled checkpoint, and 38.3% (plus or minus 0.38) for the base model. MarkTechPost reports 42.92% for the leaderboard entry and 42.24% as the multi-seed mean. These are consistent if the leaderboard number is a single submission and the others are averages over seeds. Neither BFL nor the outlets spell that out, so this reading is an inference.

Policy Parameters RoboLab-120 success
FLUX 3 Action (leaderboard entry) 7B 42.92%
FLUX 3 Action guidance-distilled (mean) 7B 42.2%
FLUX 3 Action base (mean) 7B 38.3%
Cosmos 3 Nano policy 16B 36.8%
pi0.5 3.3B 28.0%
DreamZero 14B 25.7%
GR00T N1.6 3B 7.2%

Sources: BFL model page; MarkTechPost; RoboLab paper for pi0.5 (28.0%) and GR00T N1.6 (7.2%). Baseline numbers were reported by BFL and secondary outlets, and I could not independently re-run them.

Reading the benchmark honestly

Several caveats apply. The metric is success rate over 120 tasks with 10 trials each in simulation, so a single trial moves the total by about 0.08 percentage points; a 4-point gap between guidance-distilled FLUX 3 Action and Cosmos 3 Nano is well outside the reported seed noise of about 0.4 points but says nothing about the task-distribution shift you will face. Benchmark authors choose tasks; a DROID-style Franka setup favours models fine-tuned for DROID, which is exactly where FLUX 3 Action was tuned.

Second, a 42% success rate is still a failure rate near 58%. BFL itself notes that a task such as “pumpkins in clutter” is not solved by any of the pure action policies. Third, the real-world check is small: in a blind evaluation run by Positronic Robotics on a Franka arm, FLUX 3 Action succeeded on 28 of 30 attempts (10 DROID-style tasks, 3 attempts each, 93.3%), versus 27 of 30 for Cosmos 3 Nano, 20 for DreamZero and 13 for pi0.5, per MarkTechPost. With 30 trials, the 95% interval on 28 of 30 is wide; a rough Wilson interval runs from about 79% to 98%. The ranking against Cosmos 3 Nano (28 versus 27) is not distinguishable from noise.

RoboLab-120 benchmark ablation and training pipeline for FLUX 3 Action from video pretraining to embodiment finetuning

Figure 2: The FLUX 3 Action training pipeline. Video-dominated pretraining feeds a five-source action midtraining stage, then per-embodiment fine-tuning and distillation into faster variants.

Figure 2 shows why the pipeline matters more than any single score. The fine-tuning target (DROID here, SO-101 for the other checkpoint) is a small final step. Most of the capability is decided earlier, which is why BFL ships an “action base” checkpoint with shared frozen video and text encoders for people who want to adapt to their own robot.

Latency, Hardware and Cost: Fitting a 7B World Model into a Control Loop

Reported latency

BFL publishes per-chunk inference latency on an NVIDIA B200 with FP8: 101.71 ms for the base model (4 steps, guidance off), 32.29 ms for the step-distilled one-step variant, and 182.00 ms for the guidance-distilled variant with guidance on. It also reports 21.08 ms of compute per second of robot motion on an H200. Against competitors, BFL claims 1.34x to 2.28x faster than pi0.5 across workstation and datacenter GPUs, and 3.95x faster than Cosmos 3 Nano in FP8 (9.19x to 13.22x for the step-distilled variant).

One discrepancy: DataNorth reports 41.06 ms per action block on B200 and a 2.85x to 3.15x speed-up over Cosmos 3 Nano. That does not match BFL’s table, and I could not find which configuration it refers to. Use BFL’s figures for planning, and measure on your own hardware.

Worked example: does the chunk arithmetic close?

These numbers are derived from BFL’s published figures and are illustrative for a hypothetical deployment.

  • One chunk covers 32 / 15 = 2.13 s of motion.
  • Base model on B200: 101.71 ms per chunk, so inference occupies 101.71 / 2133 = 4.8% of the chunk’s wall-clock time.
  • Step-distilled: 32.29 / 2133 = 1.5%.
  • Guidance-distilled: 182.00 / 2133 = 8.5%.

In each case the policy is much faster than the motion it commands, leaving headroom to re-plan mid-chunk. If you replan every 0.5 s (about 7 actions executed per chunk), the base model uses 101.71 / 500 = 20% of a B200’s time for one robot. A rack GPU can therefore serve several arms, but a single H200-class card at 32 GB is a large per-robot cost. The real bottleneck is often not compute but the network round trip if the GPU is remote. A 50 ms round trip on top of 100 ms inference is 150 ms of stale observation, which is fine for a 2-second chunk and bad for a reactive task like catching an object.

Sequence of camera observation, FLUX 3 Action inference, chunk buffer and controller with overlapping action chunks

Figure 3: A receding-horizon loop. The controller streams actions from a buffer while the policy computes the next overlapping chunk from a fresh observation.

Figure 3 shows the standard pattern. The controller never waits on the policy: it consumes a buffered chunk while the next one is computed. Overlap requires a blending rule where old and new chunks disagree. Blending is application logic, not part of the checkpoint, and the choice (drop the tail of the old chunk, cross-fade, or temporal ensembling) changes behaviour noticeably.

Memory and hardware

The DROID model card says the model runs in about 32 GB of GPU memory in BF16 on an H200, and fits 24 GB cards with FP8 quantisation and text-encoder offload. BFL reports testing on B200, H200, RTX 6000 Pro and RTX 5090 GPUs. A 7B model at 2 bytes per parameter is roughly 14 GB of weights in BF16; the rest of the 32 GB is activations, the video-token cache and the text encoder. The card lists Python 3.12, CUDA 12.8, PyTorch 2.10.0 and Transformers 5.16.1. The Decoder notes BFL positions it for on-device deployment, and MarkTechPost mentions NVIDIA Jetson support. I could not verify a specific Jetson module or its latency, so plan on your own benchmarks; our Jetson Thor architecture article covers the class of hardware likely to matter.

Hybrid with a reasoning model

BFL also reports a hybrid system in which a large reasoning model plans and FLUX 3 Action executes. In BFL’s account, using a model named GPT 6 Astra at low reasoning effort, the hybrid costs $8.77 and 8 minutes 8 seconds per successful task, reaches about 90% overall success, and is 29% cheaper and 40% faster than the best pi0.5 plus Astra configuration. This is the vendor’s own experiment on its own tasks, involving a third-party model I have not independently reviewed, so treat it as a demonstration of architecture (planner on top, fast policy underneath) rather than a benchmark. BFL also notes that pure-action policies still cannot solve some cluttered tasks that hybrids can.

Licence and Access: What “Open Weights” Means Here

Weights are on Hugging Face in a FLUX 3 Action collection with three repositories: flux-3-action-base (the action-adaptation base and shared frozen video and text encoders), flux-3-action-so101 (a LeRobot policy for the low-cost SO-101 arm, with a task LoRA recipe) and flux-3-action-droid (a validated DROID policy, with optimised variants in subfolders). The flux-action code repository on GitHub is Apache-2.0.

The weights fall under the FLUX Kommunity License v1.0 (the spelling is BFL’s), with the Qwen3-VL-4B text encoder under Apache-2.0. Reading the licence file in the DROID repository, the terms are:

  • Non-commercial use (research, personal study, hobby, charitable) is broadly permitted.
  • Commercial production use of outputs is limited to “Qualifying Users”, defined as entities under US$5 million annualised revenue; larger entities need a commercial licence from BFL.
  • Distillation is prohibited: you may not use the model, its intermediate representations or synthetic outputs to improve a competing model that performs a similar function.
  • Prohibited uses include military and surveillance applications, biometric processing and export-control violations, and the card says the model must not control a machine “in a way that endangers people, without human oversight”.

MarkTechPost summarises the licence as non-commercial only, which is narrower than the licence text I read. The Hugging Face collection page shows the licence field as unspecified for each repo, and DataNorth says commercial terms are unpublished. These sources do not agree, so read the LICENSE.md in the repository you download and ask BFL for anything commercial. Note also that the licence’s “Outputs” language was written for an image generator. Whether an action chunk or a fine-tuned LoRA counts as an “Output” or a “Derivative” is a question for counsel, not this article. The distillation clause deserves attention: it appears to bar using FLUX 3 Action to generate pseudo-labelled data to train your own smaller edge policy.

World Action Model vs VLA: What Actually Differs

The comparison most readers want is world action model vs VLA. Four differences matter in practice.

Where the prior comes from. A VLA inherits semantic knowledge (objects, language, common sense) from web-scale text and images. A WAM inherits dynamics knowledge (how things move and deform) from video. FLUX 3 Action’s text conditioning comes through a 4B encoder, much lighter than the language backbone in a typical VLA, so its language understanding is likely narrower. Language specificity is a known weakness in RoboLab: vague instructions degrade all policies.

What is generated. A VLA emits actions only. A WAM emits actions and imagined video. The imagined video is useful for debugging and, in principle, for verification, such as comparing a predicted frame with the next real frame to detect surprise. It also costs compute, although BFL describes the video output as optional to consume, and I have not verified how much decoding it saves.

How the action is sampled. Both families increasingly use flow matching or diffusion for actions (pi0 does), so the line is blurring. The distinguishing feature is the shared video-action backbone rather than the sampler. For the state of the art in the VLA line, see our physical AI VLA overview and pi0.5 deep dive.

Data economics. Because WAM pretraining is video-heavy, the recipe can absorb egocentric human video and game footage that VLA pipelines treat as noise. FLUX 3 Action’s 15.93% teleoperation share is small next to the VLA norm. If the ablation generalises, robot data may matter less per capability than it did in 2025.

Among WAMs, the closest peers are NVIDIA’s Cosmos 3 Nano policy (16B) and DreamZero (14B), both reported at lower RoboLab scores than FLUX 3 Action in BFL’s table. FLUX 3 Action’s distinguishing claims are size, speed and openness of weights, not a new paradigm. The paradigm is the same one analysed in the world-action-models comparison.

Trade-offs, Gotchas, and What Goes Wrong

No safety envelope. The model emits absolute joint targets with no built-in limits on velocity, force or workspace. A hallucinated chunk can command a joint jump, and the policy has no concept of a person standing in the workspace. Everything between the model and the motors is your job: joint-limit clamps, velocity and acceleration caps, self-collision and workspace checks, force or torque monitoring and a hardware emergency stop. The licence’s human-oversight clause makes this a legal issue as well as an engineering one. Our humanoid control stack article shows where such layers sit.

Simulation-to-reality gap in the evidence. The headline score is from Isaac Lab. The real-world evidence is 30 attempts on one robot type by a third party. Both point the same way, but neither covers your gripper, lighting, objects or camera placement. RoboLab itself reports strong sensitivity to wrist-camera calibration, so an off-by-a-few-degrees mount can cost more than the gap between two models.

Embodiment lock-in. The released policies target DROID-style Franka setups (three cameras, seven joints plus gripper) and the SO-101 arm. A different robot needs fine-tuning, and BFL’s own recipe (20k steps, batch size 2,048 for DROID) is a serious compute job, not a weekend LoRA. The SO-101 path with a task LoRA is lighter, but it is for a hobby-class arm.

Failure on long, cluttered, abstract tasks. RoboLab shows procedural reasoning as the hardest axis. BFL says no pure action policy solves “pumpkins in clutter”. Expect failures on multi-step tasks with ambiguous instructions, and plan a planner layer or a human fallback.

Predicted video is not truth. It is tempting to treat the imagined frames as a digital twin of the future. They are samples from a generative model and can be confidently wrong. Using them as a runtime check needs calibration against real outcomes before any safety decision depends on them.

Licence and distillation. The distillation ban blocks a common route to a smaller onboard policy. Revenue thresholds also mean growth can change your legal status.

Benchmark self-reporting. Most numbers in this post come from BFL, with baselines gathered by BFL. Independent reproductions of the RoboLab-120 result were not available at the time of writing.

Decision flow for choosing FLUX 3 Action versus a VLA or diffusion policy based on video prediction need, licence fit and hardware

Figure 4: A selection flow. Choose a world action model when predicted video helps your workflow, then test licence fit and add a safety layer regardless of model.

Decision Matrix: FLUX 3 Action vs Alternatives

The table compares the policy classes on four common needs. Ratings are my qualitative judgement from the published data, not benchmark results.

Use case FLUX 3 Action Cosmos 3 Nano policy pi0.5 (VLA) Small edge VLA or diffusion policy
Best RoboLab-120 score (reported) Highest (42.92%) 36.8% 28.0% Lower, not reported here
Research on open weights Strong, custom licence Check NVIDIA terms Open weights available Varies
Commercial product at scale Needs licence review; distillation banned Check terms Check terms Often permissive
Edge deployment 24 GB with FP8, Jetson reported Larger at 16B 3.3B, lighter Best fit
Language-rich instructions Narrower text encoder Similar caveat VLM-based, stronger prior Weak
Needs predicted video Native Native No No

Pick FLUX 3 Action when you have a Franka-like arm or SO-101, GPU headroom, a research or small-company licence fit, and value predicted video. Pick a VLA when instruction understanding dominates, or a smaller policy when the compute budget is fixed on-robot. For a landscape view including industrial adoption, see our foundation models in industrial robotics review.

Practical Recommendations

Treat FLUX 3 Action as a strong open baseline and a research substrate, not a drop-in controller. The evidence supports “best reported open model on one simulated benchmark at 7B”, not “ready for unsupervised deployment”.

Start in simulation with the same Isaac Lab and RoboLab setup, so your numbers compare with published ones. Then move to a supervised bench test with the safety layer in place before the model ever drives full-speed motion. Run the step-distilled variant first to see whether 1-step quality is enough, and move to the guidance-distilled variant if it is not; the reported gap between the two (roughly 42.2% for guidance-distilled versus 38.3% for the base model) suggests guidance matters more than step count for accuracy, though BFL does not publish a one-step RoboLab score in the material I read.

Log everything: input frames, predicted frames, action chunks and outcomes. The predicted-versus-actual frame gap is a cheap anomaly signal, and the logs feed later fine-tuning. Our LeRobot dataset v3 guide covers a storage format that fits this workflow, and the Isaac Sim, Gazebo and MuJoCo comparison helps choose a test simulator.

Checklist before a first robot trial:

  • Read LICENSE.md and confirm your revenue tier and use case.
  • Pin Python 3.12, CUDA 12.8, PyTorch 2.10.0 and Transformers 5.16.1 as on the model card.
  • Add joint, velocity and workspace clamps outside the model.
  • Fix the wrist-camera mount and log its calibration.
  • Define the chunk-blending rule and replan interval.
  • Set up a human e-stop and an operator in the loop.
  • Record a baseline (for example 30 trials per task) before changing anything.

Frequently Asked Questions

What is FLUX 3 Action?

FLUX 3 Action is a 7-billion-parameter world action model from Black Forest Labs, released September 24, 2026. It takes camera images, robot joint positions and a text instruction, and outputs both a chunk of 32 robot actions at 15 Hz (2.13 seconds) and predicted future video. It builds on the multimodal FLUX 3 backbone, which was trained mainly on video, and its weights are published on Hugging Face.

What is a world action model and how is it different from a VLA?

A world action model starts from a video generator and predicts future frames and actions together. A vision-language-action model starts from a vision-language model and predicts actions only. The WAM prior is about dynamics; the VLA prior is about semantics and language. In practice the boundary is blurring, since both use flow-matching or diffusion action heads. The shared video-action backbone is the real distinction.

What does 42.92% on RoboLab-120 mean?

RoboLab-120 is an NVIDIA-led Isaac Lab benchmark of 120 tabletop manipulation tasks on a DROID-style Franka arm. 42.92% is the success rate of FLUX 3 Action’s leaderboard entry. BFL’s multi-seed means are 42.2% for the guidance-distilled model and 38.3% for the base. It is a simulated result and vendor-reported, so it does not equal a 43% success rate on your robot.

Is FLUX 3 Action free for commercial use?

Not automatically. It uses the FLUX Kommunity License v1.0. The licence text I read permits commercial production use only for entities under US$5 million annualised revenue and requires a commercial licence otherwise, and it forbids using the model to distil competing models. Some outlets describe it as non-commercial only, so check the LICENSE.md in the repository and contact BFL for commercial terms.

What hardware do I need to run FLUX 3 Action?

The model card says about 32 GB of GPU memory in BF16 on an H200, or a 24 GB card using FP8 quantisation with text-encoder offload. BFL tested on B200, H200, RTX 6000 Pro and RTX 5090. On a B200 with FP8, the base model takes 101.71 ms per 2.13-second chunk, and the one-step variant 32.29 ms. Embedded Jetson support is reported but I could not verify specifics.

Is it safe to connect FLUX 3 Action directly to a robot?

No. The model outputs absolute joint targets and has no built-in velocity, force or workspace limits, as BFL states. The licence also disallows controlling machines in ways that endanger people without human oversight. Put clamps, collision and workspace checks, force monitoring and a hardware emergency stop between the policy and the motors, and keep an operator supervising until you have logged extensive trials.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *