ROS 2 Nav2 in Production: How Autonomous Mobile Robots Navigate Complex Warehouse Environments

ROS 2 Nav2 in Production: How Autonomous Mobile Robots Navigate Complex Warehouse Environments

Nav2 Autonomous Mobile Robot Navigation in Warehouses: A Production Guide

A warehouse is a hostile place for a robot that thinks in straight lines. Pallets appear in aisles that were empty an hour ago, forklifts cross at blind corners, shrink-wrap glints back at the LiDAR as a phantom wall, and a hundred other robots want the same intersection. Getting a single demo to drive from A to B is a weekend project. Keeping a fleet moving through a shift without a human walking over to rescue a stuck unit is an engineering discipline.

The most widely adopted open-source answer to that problem is Nav2, the ROS 2 navigation framework. A Nav2 autonomous mobile robot is not one algorithm. It is a set of lifecycle-managed servers, a behavior tree that sequences them, layered costmaps that turn sensor noise into a cost surface, and a safety layer that can override all of it. This guide explains how those pieces fit, what to tune for warehouse work, how fleets sit on top, and where deployments actually break.

What this covers: the Nav2 server architecture and why it is shaped that way, the default behavior tree and its recovery logic, costmaps and localization, the planner and controller choices (including MPPI versus Regulated Pure Pursuit), concrete YAML for a warehouse profile, fleet coordination with Open-RMF and VDA5050, and the failure modes worth designing against.

Context and Background

Nav2 is the successor to the ROS 1 move_base navigation stack. Its design was published at IROS 2020 in “The Marathon 2: A Navigation System” by Steve Macenski, Francisco Martin, Ruffin White and Jonatan Gines Clavero. The paper’s headline changes were a behavior tree for navigator task orchestration, methods suited to dynamic and populated environments, and a foundation on ROS 2 with safety-critical use in mind. The authors validated it by running a robot alongside students on a campus over a marathon-length distance. You can read the original on arXiv.

Since then the project has matured into a community-governed stack with professional maintenance provided through Open Navigation LLC, and its repository lists sponsors including Dexory, NVIDIA, AMD, Polymath Robotics, Stereolabs and 3Laws Robotics. The project’s build matrix covers the Humble, Jazzy and Lyrical ROS 2 distributions. The Lyrical Luth ROS 2 release shipped in May 2026, and the Nav2 team published its Lyrical feature summary on 25 August 2026. If you are weighing a distribution move, our ROS 2 Kilted to Lyrical Luth migration guide covers the platform side, and the Nav2 Lyrical versus Kilted comparison goes deeper on the controller changes.

Older write-ups, including an earlier version of this article, repeat a claim that Nav2 is “trusted by over 100 companies.” I could not verify that figure against a primary source, so this rewrite does not rely on it. What can be verified is the sponsor list above and the fact that Nav2 is the default navigation option in the ROS 2 ecosystem that most AMR start-ups and several large vendors build on.

Why does that matter for a warehouse? Because the commercial alternatives, vendor-locked navigation stacks bundled with a robot base, trade away configurability for convenience. Nav2 gives you the planner, the controller, the costmap layers and the recovery logic as plugins you can inspect, replace and test. The price is that the integrator owns the tuning, and the defaults are tuned for a small research platform, not a 600 kg pallet mover.

Three facts frame everything that follows. First, Nav2 is deliberately modular: each task (plan, smooth, control, recover) lives in its own action server so that algorithms can be swapped through pluginlib without recompiling the stack. Second, the glue is a behavior tree, which makes the recovery policy data rather than code. Third, Nav2 assumes someone else provides localization and a map: it consumes the transform tree and a map, and it does not care whether those come from AMCL, SLAM Toolbox, or a commercial localizer. We cover the mapping side in more depth in our SLAM architecture guide.

The Nav2 Reference Architecture for a Warehouse AMR

In short: Nav2 runs four action servers (planner, smoother, controller, behavior) under a BT Navigator that executes a behavior tree. Costmaps feed the planner and controller, a transform tree supplies the robot pose, and a Collision Monitor can veto velocity commands before they reach the base. A fleet manager sits above, sending goals.

Nav2 autonomous mobile robot reference architecture showing servers, costmaps, collision monitor and fleet manager

Figure 1: Nav2 autonomous mobile robot reference architecture. The fleet manager sends goals to the BT Navigator, which orchestrates the planner, smoother, controller and behavior servers. The Collision Monitor sits between the controller output and the robot base.

Figure 1 shows the dependency structure. The fleet manager (or a human using RViz) issues a NavigateToPose action goal. The BT Navigator loads a behavior tree XML and ticks it. The tree calls into the servers, each of which hosts algorithm plugins. Two costmaps, global and local, are built from the map and live sensor data. The controller publishes velocity commands, which in a production stack pass through the Collision Monitor and, optionally, a velocity smoother before the base driver sees them.

Servers, plugins and lifecycle management

The Nav2 documentation describes four primary action servers. The Planner Server computes a valid and ideally optimal path from the current pose to a goal. The Controller Server executes path following using local environmental data and produces the control effort for the base. The Smoother Server refines paths by reducing raggedness and increasing obstacle clearance. The Behavior Server handles recovery behaviors and fault tolerance, each with its own action interface.

Each server hosts algorithms as pluginlib plugins, which means a configuration line selects Smac Hybrid-A instead of NavFn, or MPPI instead of Regulated Pure Pursuit, without touching code. The documented plugin catalogue is broad. Planners include NavFn, Smac 2D, Smac Hybrid-A, the State Lattice planner and Theta*. Controllers include DWB, TEB, Regulated Pure Pursuit, MPPI, a Rotation Shim controller, a Pose Following controller and a Vector Field controller. Smoothers include a Simple Smoother, a Constrained Smoother and a Savitzky-Golay Smoother.

Every server is a ROS 2 lifecycle node, and that matters operationally. Lifecycle nodes move deterministically through unconfigured, configured, active, deactivated and finalized states. Nav2 wraps them in its own lifecycle node class that adds a bond connection, so the lifecycle manager notices if a server dies. In a fleet this is how you get a robot that refuses to accept goals until every server is healthy, instead of one that accepts a goal and silently never moves.

The transform tree and the frames Nav2 expects

Nav2 follows ROS REP-105: map to odom to base_link to the sensor frames. The localization system publishes the map to odom transform, which is allowed to jump when a correction arrives. The odometry source publishes odom to base_link, which must be continuous and smooth. This split is the single most important idea in the stack, and violating it is the root of many “robot lurches at the aisle end” bug reports.

The reason is practical. The local controller needs a pose that never teleports, otherwise a trajectory that was safe one tick ago becomes nonsense the next. The global planner, by contrast, wants the best estimate of where the robot sits on the map. Splitting the two frames lets each consumer pick the one it needs. If you feed a wheel odometry source that glitches (for example, a slipping wheel reporting a velocity spike), the local controller reacts to it as if it were real, so odometry quality deserves as much attention as the LiDAR.

Why a 2D costmap is still the workhorse

Nav2 represents the environment as what the docs call a regular 2D grid of cells containing a cost: unknown, free, occupied, or inflated. Layers are plugins. The documented set includes a Static Layer (the map), an Obstacle Layer (persistent 2D costmap from laser scans with raycasting), a Voxel Layer (persistent 3D voxel data from depth and laser readings), a Range Sensor Layer, an Inflation Layer (inflates lethal obstacles with exponential decay), a Spatio-Temporal Voxel Layer with decay, and an Obstacle Smoothing Layer that filters noise-induced standalone obstacles.

For warehouses, a 2D grid is usually enough because the robot is a ground vehicle operating on a flat slab. The problems a 2D costmap cannot see, such as a pallet fork tine sticking out at 40 cm or an overhanging shelf at 1.2 m, are better solved with a depth camera feeding the Voxel Layer or with a second, higher-mounted LiDAR, rather than by moving to a full 3D planner. That is a design opinion, but it follows from the cost: a 3D representation multiplies memory and compute for every planner and controller tick.

Two costmap instances run in parallel. The global costmap covers the whole map at lower update rates and feeds the planner. The local costmap is a rolling window centered on the robot, updated faster, and feeds the controller. The costmap_2d defaults in the Nav2 configuration reference are an update_frequency of 5.0 Hz, a publish_frequency of 1.0 Hz and a resolution of 0.1 m per cell, with a default robot_radius of 0.1 m. Those are starting points. A warehouse AMR with a 0.8 m wide footprint should define a polygon footprint and not rely on the radius default.

The Behavior Tree: Where Navigation Policy Lives

In short: the BT Navigator ticks a behavior tree (BehaviorTree.CPP V4) that decides when to plan, when to follow, and what to try when something fails. The shipped default replans at a fixed rate and escalates through clearing costmaps, spinning, waiting and backing up. Because it is XML, you can change recovery policy per site without rebuilding.

Nav2 behavior tree with replanning and round-robin recovery nodes

Figure 2: Simplified Nav2 behavior tree for navigate-to-pose with replanning and recovery. A PipelineSequence runs the planning branch and the following branch; failures escalate to a round-robin of recovery actions.

Nav2 uses BehaviorTree.CPP V4 and the documentation argues, convincingly, that trees are more maintainable than finite state machines for multi-step autonomy because primitives are reusable and composable. Our behavior trees for robot task planning article covers the general theory. Here we stay with how Nav2 uses them.

Control nodes that make Nav2 trees work

Two Nav2-specific control nodes are worth understanding because they define the runtime behavior. The PipelineSequence ticks its first child until it succeeds, then ticks the first and second children until the second succeeds, and so on. A RUNNING child does not change the flow, and a FAILURE from any child halts the sequence. This is how Nav2 keeps the planner ticking while the controller follows the latest path: planning is the first child, following is the second, and both run concurrently.

The RecoveryNode wraps an action and a recovery. If the main child fails, the recovery child runs and the main child is retried, up to a configured retry count. The RoundRobin node cycles through its children across successive failures so that a robot that failed once by spinning tries waiting next time. The docs also list a PersistentSequence, a NonblockingSequence and, new in Lyrical, a PauseResumeController node.

Action nodes include ComputePathToPose, SmoothPath, FollowPath, Spin, Wait, BackUp, DriveOnHeading, AssistedTeleop, ClearCostmap services, ReinitializeGlobalLocalization, DockRobot and UndockRobot. Conditions include goal reached, initial pose set, progress checks, battery checks, transform availability, path validity and timers. Decorators throttle by rate, distance or speed, and notify on goal updates.

The default navigate-to-pose tree, step by step

The tree shipped with Nav2 (based on my reading of the repository’s default XML; check the file in your distribution since it has been evolving) follows this pattern. A RateController wraps path computation so the planner replans at about 1 Hz rather than every tick. The planning branch is a RecoveryNode: if ComputePathToPose fails, clear the global costmap and retry. The following branch is a second RecoveryNode: if FollowPath fails, clear the local costmap and retry. If both of those local retries fail, control passes to a top-level recovery branch, a RoundRobin of clear-both-costmaps, Spin, Wait and BackUp.

That ordering is deliberate. Clearing costmaps fixes the common perception failure where a transient object left a lethal ghost. Spin gives the sensors a fresh look and may open a path. Wait handles time-based obstacles such as a forklift crossing. BackUp gets the robot out of a tight corner. Each is cheap and reversible, and they escalate from least to most physically committal.

Customizing the tree for warehouse operations

A warehouse deployment almost always edits the tree. Three modifications recur in practice (these are patterns, not a Nav2 requirement). First, make Wait longer and prefer it before Spin in aisles: a 360 degree spin in a 1.4 m aisle with a wide robot is a collision waiting to happen, and an operator watching it will lose faith in the system. Second, cap the total recovery time and then report failure upward to the fleet manager, so the task can be reassigned to another robot rather than retried for ten minutes. Third, insert a condition node for battery level so the robot refuses new long goals when it should dock.

The Lyrical release adds a PauseResumeController node, which gives a clean way to pause a navigation task on an external trigger and resume it, which maps naturally to a human-crossing zone or a fleet-level hold. Previously integrators approximated this by cancelling and re-sending goals, which loses progress state.

Behavior trees also give you an audit trail. Because each node returns SUCCESS, FAILURE or RUNNING and the tree is logged, you can reconstruct why a robot did what it did. For incident review in a regulated facility that is worth more than any benchmark number.

Localization, Planning and Control in Depth

In short: for a mapped warehouse, run AMCL or SLAM Toolbox in localization mode for the map to odom transform, a cost-aware global planner (Smac for car-like or large robots, NavFn or Smac 2D for small holonomic ones) and pick between Regulated Pure Pursuit for predictable industrial motion or MPPI for dynamic, cluttered spaces.

Nav2 goal execution sequence from fleet manager through planner, controller and collision monitor to the robot base

Figure 3: Sequence of a single Nav2 goal. The planner produces a path once and then at the replanning rate, the controller closes the loop every tick, and the Collision Monitor gates the final velocity command.

Localization: AMCL versus SLAM Toolbox in production

AMCL, the adaptive Monte Carlo localizer, keeps a particle set representing pose hypotheses and weights them against a laser scan and a static map. The documented defaults are 500 to 2,000 particles, a likelihood_field laser model, updates triggered after 0.25 m of translation or 0.2 rad of rotation, and odometry noise parameters alpha1 to alpha5 all at 0.2, with a differential-drive motion model. These defaults suit a modest indoor robot. A warehouse changes the problem in two ways: long featureless aisles and repeated geometry.

Long aisles are the classic failure. Two parallel racks look the same along the aisle axis, so the filter cannot tell whether the robot has moved 2 m or 4 m down the aisle until an end feature appears. The filter’s weights drift along that axis, and the pose slowly slides. Mitigations that practitioners use include adding reflectors or distinctive fiducials at aisle ends, fusing wheel and IMU odometry in an extended Kalman filter so that the dead-reckoning prior is strong, and lowering update_min_d so the filter integrates more frequently. Treat those as engineering mitigations rather than guaranteed fixes.

SLAM Toolbox is the other option. It supports lifelong mapping and a localization mode that works against a serialized pose graph. In a warehouse whose racking changes weekly, running localization against a map that is periodically refreshed, rather than a frozen one, can reduce the “map no longer matches reality” failures. The cost is operational: you now manage map versions, validation before rollout, and rollback. Our SLAM architecture article walks through the pose graph mechanics and loop closure trade-offs, and the multi-sensor fusion guide covers the odometry fusion that sits underneath either choice.

One more production rule: never let the robot start without a pose check. The BT includes an initial-pose-set condition for a reason. A fleet that auto-starts robots in the wrong place after a power cycle produces the most expensive class of incident, a confident robot that is wrong.

Global planners: what each is good for

The global planner runs rarely (on the order of once per second in the default tree) and produces a path through the global costmap. The choice depends on the robot kinematics.

NavFn uses A or Dijkstra expansion and assumes a 2D holonomic particle. It is fast and fine for small, round differential-drive robots. Smac 2D is a 2D A with 4- or 8-connected neighborhoods. The Smac Hybrid-A* planner is an SE2 implementation using Dubins or Reeds-Shepp motion models, which means it plans paths the vehicle can actually drive, honoring a minimum turning radius. Its documented defaults include a minimum_turning_radius of 0.4 m, DUBIN motion model, 72 angle quantization bins, max_iterations of 1,000,000 and max_planning_time of 5.0 s. The State Lattice planner uses pre-generated minimum control sets and suits non-circular or non-holonomic bases that need precise primitives.

For a warehouse with a rectangular, differential-drive base that turns in place, Smac 2D or NavFn plus a smoother is often adequate and cheap. For an Ackermann or tugger-style vehicle that cannot turn in place, Hybrid-A* or State Lattice is the correct tool, because a 2D plan will be undrivable and the controller will thrash. The honest rule: pick the planner whose motion model matches the vehicle, then spend your tuning time on the costmap, where most real failures live.

Paths from 2D planners are jagged, so the Smoother Server matters. The Simple Smoother is a lightweight option, the Constrained Smoother optimizes criteria with a solver, and the Savitzky-Golay Smoother applies a digital-signal-processing filter. A smoother also increases clearance from obstacles, which is the difference between a robot that scrapes racking corners and one that does not.

Controllers: Regulated Pure Pursuit versus MPPI

The controller is where Nav2 deployments diverge most. The two controllers that matter most for AMRs are Regulated Pure Pursuit (RPP) and Model Predictive Path Integral (MPPI). The older DWB (an implementation of the Dynamic Window Approach with plugin interfaces) and TEB (an MPC-like controller) remain available.

Regulated Pure Pursuit is described in the docs as a variation on pure pursuit targeting service and industrial robots. It follows the path with a lookahead point and adds regulation: it slows in proportion to path curvature, slows near obstacles, and scales its lookahead with speed. Defaults include a desired_linear_vel of 0.5 m/s, lookahead_dist of 0.6 m (with min 0.3 m and max 0.9 m when velocity-scaled), regulated_linear_scaling_min_radius of 0.90 m, regulated_linear_scaling_min_speed of 0.25 m/s and cost_scaling_dist of 0.6 m, with collision detection enabled. Its strengths are predictability and low compute: the robot follows the planned path closely, and when you see odd behavior it is easy to reason about why.

MPPI is a sampling-based predictive controller. It samples many candidate control sequences, rolls them out through a motion model, scores them with a set of critics, and takes a weighted combination. The documented defaults are 1,000 sampled trajectories (batch_size), 56 time steps of 0.05 s each (a 2.8 s prediction horizon), a vx_max of 0.5 m/s and a DiffDrive motion model. The default critics include Constraint, Cost, Goal, GoalAngle, PathAlign, PathFollow, PathAngle and PreferForward, with Obstacles, VelocityDeadband and Twirling available. MPPI handles dynamic obstacles and omnidirectional or Ackermann models gracefully and tends to produce smooth, human-like trajectories. The cost is CPU time and a bigger tuning surface, since each critic has weights.

The Lyrical release materially extends MPPI. According to the Nav2 team’s summary, it adds an open-loop control mode, dynamics support, asymmetric acceleration limits, plugin motion models, trajectory validators and delay compensation. It also adds Dynamic Window Pure Pursuit to the RPP controller, to account for dynamic limits during tracking, and an Adaptive Goal Checker with coarse and fine criteria. We compare these in our Lyrical versus Kilted analysis.

A decision matrix

Situation Planner Controller Reason
Round differential-drive AMR, mostly static aisles NavFn or Smac 2D + smoother RPP Cheap, predictable, easy to certify behavior
Rectangular AMR, tight aisles, frequent pedestrians Smac 2D + smoother MPPI Samples evasive trajectories, handles dynamic obstacles
Tugger or Ackermann vehicle Smac Hybrid-A* or State Lattice MPPI with Ackermann model or RPP Kinematically feasible plans required
Omnidirectional base Smac 2D MPPI with Omni model Exploits lateral motion
Very low-power compute (small ARM board) NavFn RPP Lowest CPU load

The pairings above are my engineering judgment based on the documented controller properties, not a vendor benchmark.

Rotation Shim and collision-aware motion

A frequent warehouse pain point is the robot driving off at the wrong heading when a new path starts behind it. The Rotation Shim Controller solves this: it rotates to the path heading before passing control to the main controller. Pair it with RPP so that the pure pursuit logic never has to cope with a large initial heading error. It is a small configuration change that removes a visible class of weird starts.

Warehouse Tuning: Concrete Configuration

In short: tune the footprint and inflation first, size the local costmap to your stopping distance, then match controller speed limits to what the safety scanners are certified to stop. Most “Nav2 is flaky” complaints trace to footprint, inflation, odometry or frame configuration, not the algorithms.

The following YAML is an illustrative starting profile for a differential-drive pallet-carrying AMR, not a validated production file. Numbers in comments are choices you should replace after measurement. Parameter names follow the Nav2 configuration reference; the specific values are mine.

controller_server:
  ros__parameters:
    controller_frequency: 20.0
    controller_plugins: ["FollowPath"]
    progress_checker_plugins: ["progress_checker"]
    goal_checker_plugins: ["general_goal_checker"]
    FollowPath:
      plugin: "nav2_regulated_pure_pursuit_controller::RegulatedPurePursuitController"
      desired_linear_vel: 1.0          # choose from safety scanner stopping data
      use_velocity_scaled_lookahead_dist: true
      min_lookahead_dist: 0.5
      max_lookahead_dist: 1.2
      use_regulated_linear_velocity_scaling: true
      regulated_linear_scaling_min_radius: 1.2
      regulated_linear_scaling_min_speed: 0.2
      use_cost_regulated_linear_velocity_scaling: true
      use_collision_detection: true
      allow_reversing: false

local_costmap:
  local_costmap:
    ros__parameters:
      update_frequency: 10.0
      publish_frequency: 5.0
      rolling_window: true
      width: 6
      height: 6
      resolution: 0.05
      footprint: "[[0.6, 0.4], [0.6, -0.4], [-0.6, -0.4], [-0.6, 0.4]]"
      plugins: ["obstacle_layer", "inflation_layer"]
      inflation_layer:
        plugin: "nav2_costmap_2d::InflationLayer"
        cost_scaling_factor: 3.0
        inflation_radius: 0.7
      obstacle_layer:
        plugin: "nav2_costmap_2d::ObstacleLayer"
        observation_sources: scan
        scan:
          topic: /scan
          data_type: "LaserScan"
          marking: true
          clearing: true
          obstacle_max_range: 5.0
          raytrace_max_range: 6.0

Footprint and inflation: the two settings that decide everything

Replace robot_radius with a measured polygon footprint that includes any load overhang, then set inflation so the robot prefers the middle of an aisle. The inflation layer assigns lethal cost inside the footprint, then a decaying cost out to inflation_radius. A larger cost_scaling_factor makes cost fall off faster, so the robot is willing to run closer to racking. The default cost_scaling_factor of 1.0 is a gentle gradient that, with a wide robot, can make the planner refuse narrow but passable aisles because inflated cost swallows the gap.

Check that inflation_radius is not so large that two inflated racking faces overlap across a 1.5 m aisle. When they overlap, every cell in the aisle carries high cost, the planner treats the aisle as nearly blocked, and you get “no valid path” errors for routes a human would call trivial. Visualize the inflated costmap in RViz before you blame the planner. Lyrical’s inflation layer work, which the team describes as massive speed-ups through new algorithms and optional OpenMP, makes large inflation radii cheaper to compute, but does not change the geometry.

Sensor ranges, resolution and update rates

Set obstacle_max_range shorter than the sensor’s trustworthy range, and raytrace_max_range slightly longer so clearing reaches beyond marking. A 0.05 m resolution doubles linear detail over the 0.1 m default and quadruples cell count, so keep the local window small (6 m by 6 m above) and the global costmap at coarser resolution. As a worked example, a 6 m by 6 m window at 0.05 m resolution is 120 by 120, or 14,400 cells, which is trivial to update at 10 Hz. A 100 m by 80 m global map at the same resolution is 2,000 by 1,600, or 3.2 million cells, which is why global costmaps update slowly.

Controller frequency should match the control loop the base can actually track. Setting 50 Hz on a base that applies commands at 20 Hz wastes CPU and increases jitter. A velocity smoother between the controller and the base then limits acceleration and jerk, which protects loads and wheels.

Match speed to stopping distance

The most important safety number is not a Nav2 parameter at all. It is the stopping distance of the physical robot at its maximum speed with its heaviest load, which the safety scanner and brake design must satisfy. The Nav2 desired_linear_vel should be derived from it, with margin. Nav2 is not a safety-rated system by itself; certified protective stops must come from a safety-rated scanner and safety controller, with Nav2’s Collision Monitor providing the software layer that avoids most trips in the first place.

Safety Layers and Fleet Coordination

In short: Nav2 plans for one robot. Warehouses need a layer above that arbitrates shared space (Open-RMF or a VDA5050 master control) and a layer below that stops the robot when planning fails (the Collision Monitor and a safety-rated scanner).

Fleet architecture with dispatcher, traffic scheduler, fleet adapters, Nav2 on each AMR and safety layers

Figure 4: Fleet-level architecture. A task dispatcher and traffic scheduler assign work and reserve space through fleet adapters, each robot runs its own Nav2 instance, and local safety layers operate independently of the fleet.

The Collision Monitor: a last software line of defense

Nav2’s Collision Monitor performs collision avoidance directly from sensor data, bypassing the costmap and trajectory planners, to prevent collisions at an emergency level. The companion Collision Detector is passive: it only publishes alerts. The documented action types are stop, slowdown, limit and approach, and it supports multiple polygon types and data sources. Lyrical adds toggles, exclusion zones and costmap sources.

Why bypass the costmap? Because the costmap pipeline has latency (sensor, raycasting, inflation, planner, controller) and can be wrong in ways that persist. A polygon check on the most recent scan, applied to the final cmd_vel, has far fewer ways to fail. A common configuration is a slowdown polygon that scales speed when something enters a zone ahead of the robot and a smaller stop polygon for emergencies, with the polygon dimensions tied to the velocity so the stop zone grows with speed.

Be precise about what this is. The Collision Monitor is software running on a general-purpose computer. It reduces the number of safety-scanner trips and bumps, and it improves behavior, but it is not a substitute for a safety-rated laser scanner on a safety PLC meeting the applicable AMR safety standards. Design with both.

Why one robot’s Nav2 is not a fleet

Two robots each running a correct Nav2 stack will still deadlock head-on in a one-lane aisle, because each treats the other as a dynamic obstacle and each waits, or each reverses, in symmetrical lockstep. Local obstacle avoidance cannot resolve a negotiation that requires a global decision about who yields and where. This is the central argument for a fleet layer, and the cleanest way to see it is that Nav2 optimizes the path of a single agent while a fleet needs a schedule of space and time.

Open-RMF (Robotics Middleware Framework) is described by its maintainers as a free, open source, modular system that enables sharing and interoperability between multiple fleets of robots and physical infrastructure such as doors, elevators and building management systems. Since 2024 it is governed under the Open Source Robotics Alliance. It was originally developed with healthcare in mind and is positioned for factories and distribution centers as well. Architecturally the idea is that each robot vendor’s fleet is wrapped by a fleet adapter, a traffic scheduler reserves lane and time slots across fleets, and infrastructure adapters let robots request doors and lifts. We walk through a hands-on setup in our Open-RMF fleet orchestration tutorial.

For homogeneous fleets in conventional warehouses, another route is the VDA5050 interface, an open standard that defines the messages exchanged between a master control and automated guided vehicles over MQTT. It standardizes orders, state and instant actions rather than motion, so the robot’s own navigation stays in charge of how to drive between nodes. Our VDA5050 AMR fleet management article covers the message flow. In practice the integration pattern is a thin bridge node that translates a VDA5050 order (a graph of nodes and edges) into a sequence of Nav2 NavigateToPose or follow-waypoints goals and maps Nav2 results back into VDA5050 state.

The two approaches differ in emphasis. VDA5050 gives you a vendor-neutral protocol between a fleet controller and heterogeneous vehicles. Open-RMF gives you the traffic negotiation and infrastructure integration. They are complementary more than competing, and a sensible reading is: pick the protocol for how the fleet talks, and pick the scheduler for how it shares space. The exact integration between the two is deployment-specific, and I have not verified a canonical adapter, so evaluate the current community packages before committing.

Lane graphs, not open floor, at fleet scale

A warehouse at scale often constrains robots to a lane graph, a directed graph of waypoints and lanes with one-way and reserved segments. Nav2 still handles the local geometry between waypoints (avoiding a dropped box) while the fleet layer reserves the lane. The trade-off is flexibility. A lane graph is easy to reason about, to simulate and to certify, but less efficient than free navigation, and it requires editing when racking moves. The Nav2 team’s Route Server work, which targets graph-based routing inside Nav2, may reduce the gap between those worlds; I did not verify its status for the Lyrical release, so check the current docs.

Simulate before you drive

Never test fleet behavior first on real racking. A simulator with the real map, real footprint and a traffic model lets you replay the bad shift. Our comparison of Isaac Sim, Gazebo and MuJoCo outlines the simulator trade-offs. The goal is a regression suite: record a rosbag from a failure, replay it, change one parameter and compare outcomes.

Trade-offs, Gotchas, and What Goes Wrong

In short: failures cluster in perception artifacts, localization drift in repetitive geometry, parameter coupling between planner, controller and costmap, and recovery behaviors that are unsafe in tight spaces.

Phantom obstacles. Reflective shrink-wrap, glossy floors and glass can produce spurious LiDAR returns that mark as lethal obstacles. The obstacle layer raycasts to clear them, but only if the clearing ray passes through the phantom cell. The Obstacle Smoothing Layer filters standalone noise-induced obstacles, which is a documented tool for this. Costmap clearing recoveries help but treat the symptom.

Localization drift in aisles. Covered above. The signature is a robot that is fine in open areas and wanders sideways in long racks. Detect it by comparing the scan to the map: if scan-to-map alignment quality (AMCL’s covariance, or a custom metric) degrades, stop and relocalize rather than continuing.

Parameter coupling. The inflation radius changes planner paths, which changes how close the controller drives to cost, which changes how often the Collision Monitor fires. Tune in order: footprint, inflation, planner, smoother, controller, collision monitor. Changing one parameter and then evaluating on a single run is a common error; evaluate on a recorded set of routes.

Oscillation at goals. A goal checker that is too strict relative to the controller’s tracking accuracy makes the robot creep back and forth around the goal. The Lyrical Adaptive Goal Checker, with coarse and fine criteria, addresses part of this; otherwise loosen tolerance or use a pre-docking approach with a separate precision mode.

Recovery in tight spaces. The default Spin and BackUp behaviors assume free space. In a 1.4 m aisle with a wide robot, a spin collides. Remove or constrain recoveries (for example drive-on-heading or wait only) and report failure to the fleet manager early.

Dynamic humans. Nav2 treats people as moving obstacles in the costmap. It does not predict intent. A controller like MPPI can sample evasive motions, but nothing replaces speed limits and safety zones in mixed human-robot areas.

Compute and middleware load. Running perception, localization, costmaps and a controller on one small board can starve the controller loop. Lyrical’s inter-process communication support for zero-copy transport in composed bringups is aimed at that, and ROS 2 Lyrical keeps Fast DDS as the default middleware with Native Buffers support in the Fast DDS RMW. Measure controller loop jitter under load before shipping, and do not assume defaults will hold.

What I could not verify. No public, apples-to-apples benchmark exists that ranks RPP against MPPI across warehouse layouts, so any claim that one is “faster” or “safer” in general is anecdotal. Treat the matrix above as guidance and measure on your own routes.

Practical Recommendations

If you are starting a warehouse AMR on Nav2 today, begin with a stack you can reason about, and add sophistication only when logs show you need it. Use the current distribution that your hardware vendor supports, and plan the move to Lyrical deliberately, since its feature set (MPPI improvements, pause and resume, adaptive goal checking, faster inflation) addresses several warehouse pain points.

Start with RPP plus a Rotation Shim and a simple smoother. Move to MPPI when pedestrians or clutter make path-following brittle. Keep recoveries conservative. Put the Collision Monitor in the chain, with a hardware safety scanner under it. Add a fleet layer when you have more than a handful of robots sharing aisles, and test deadlock scenarios in simulation first.

  • Measure the real footprint and replace robot_radius with a polygon.
  • Derive maximum speed from measured stopping distance, then set desired_linear_vel.
  • Visualize inflated costmaps in RViz and confirm every aisle has passable cost.
  • Fuse wheel odometry with an IMU and verify the odom frame is smooth.
  • Add fiducials or reflectors where aisles are featureless.
  • Customize the behavior tree: remove unsafe spins, cap total recovery time, report upward.
  • Log behavior tree transitions and record rosbags for every failure.
  • Replay failures in simulation before changing parameters in production.
  • Version your maps and parameters together and keep a rollback path.
  • Treat Nav2 as the navigation layer, not the safety layer.

Frequently Asked Questions

What is Nav2 and how does it work on an autonomous mobile robot?

Nav2 is the ROS 2 Navigation Framework. It runs planner, smoother, controller and behavior action servers under a BT Navigator that executes a behavior tree. Given a goal pose, it computes a path on a global costmap, tracks it with a controller using a local costmap, and triggers recovery actions on failure. It relies on external localization to supply the map to odom transform.

Is Nav2 production ready for warehouse robots?

Yes, with qualifications. Nav2 has professional maintenance, corporate sponsors and a long public track record, and it is used as a base by many AMR developers. But production readiness is a property of your integration: footprint, tuning, sensing, safety certification and fleet orchestration are yours to deliver. Nav2 is not itself a safety-rated product and must be paired with certified safety hardware.

Should I use MPPI or Regulated Pure Pursuit?

Use Regulated Pure Pursuit when you want predictable, low-compute path following in structured aisles. Use MPPI when dynamic obstacles, clutter or non-trivial kinematics require sampling evasive trajectories, and you can afford the CPU and tuning effort. Lyrical improves both: MPPI gains trajectory validators and open-loop mode, and RPP gains Dynamic Window Pure Pursuit. Test both on recorded routes from your site.

Which Nav2 planner is best for warehouses?

It depends on kinematics. For small differential-drive robots, NavFn or Smac 2D with a smoother is usually enough. For vehicles that cannot turn in place, such as tuggers or Ackermann platforms, Smac Hybrid-A* or the State Lattice planner generate feasible paths. The planner matters less than correct footprint and inflation settings, which drive most path-quality problems.

How do you manage a fleet of Nav2 robots?

Nav2 handles single-robot navigation, so fleets add a coordination layer. Open-RMF provides a traffic scheduler, fleet adapters and infrastructure integration, while VDA5050 standardizes the order and state messages between a master control and vehicles. A bridge translates fleet orders into Nav2 goals and returns the state. Design lanes and test deadlock cases in simulation first.

How is Nav2 different from the ROS 1 navigation stack?

Nav2 replaces the monolithic move_base state machine with modular action servers, pluginlib algorithm plugins and a behavior tree for orchestration, all on ROS 2 lifecycle nodes and DDS middleware. That makes recovery policy configurable data, adds modern planners such as Smac and controllers such as MPPI, and supports deterministic startup and health monitoring through lifecycle bonds.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *