PAPER REPORTENAll readings ↗

FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Chunran Zheng; Wei Xu; Zuhao Zou; Tong Hua; Chongjian Yuan; Dongjiao He; Bingyang Zhou; Zheng Liu; Jiarong Lin; Fangcheng Zhu; Yunfan Ren; Rong Wang; Fanle Meng; Fu Zhang

Affiliations: Mechatronics and Robotic Systems (MaRS) Laboratory, Department of Mechanical Engineering, University of Hong Kong, Hong Kong SAR, China; Information Science Academy of China Electronics Technology Group Corporation

Source: IEEE Transactions on Robotics · ref-ae98866e6b155b0f253c ↗ · Catalog record

Reading: 391 / 558 · 6 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: FAST-LIVO2 uses LiDAR geometry to stabilize sparse visual alignment in a shared-map filter, trading dependence on calibrated sensors and local photometric assumptions for efficient odometry. e-identitye-probleme-overviewe-filtere-benchmark

At a glanceWhat to know
Research problem
Source description

LiDAR loses geometric constraints near single planes, while cameras suffer from weak texture, blur and changing illumination. FAST-LIVO2 seeks complementary constraints without expensive feature matching or separate visual depth reconstruction, under the computational limits of onboard robotics. e-probleme-overview

Core mechanism
Source description

A sequential ESIKF fuses heterogeneous measurements while allowing separate visual pyramid updates. One adaptive voxel map supplies both LiDAR planes and visual patch anchors. e-overviewe-filtere-map

A key reported resultPublic-dataset odometry: Default: 0.045; with normal refinement: 0.044.

Absolute translational RMSE (m), reported Average row. 25 NTU-VIRAL, Hilti’22 and Hilti’23 sequences; selected single cameras at 10 Hz; Hilti’23 Site 3 excluded. LVI-SAM loop closure disabled.

FAST-LIVO: 0.137; FAST-LIO2: 0.151; R3LIVE: 0.278. Table II separates the two FAST-LIVO2 configurations; adjacent prose highlights 0.044. These reported averages have no repeat-run uncertainty. Baselines include adapted implementations, and failures are marked separately. e-datasetse-confige-benchmark

Reading caution
Source description

Long-distance drift remains because loop closure and sliding-window optimization are future work. Accurate timing/extrinsics are assumed. Direct alignment still needs a useful prior; degraded images can reduce accuracy, and exposure modeling omits camera response and vignetting. e-conclusione-statee-probleme-ablatione-exposure

Core contributions

  • Source description

    A sequential ESIKF fuses heterogeneous measurements while allowing separate visual pyramid updates. One adaptive voxel map supplies both LiDAR planes and visual patch anchors. e-overviewe-filtere-map

  • Source description

    LiDAR plane priors guide patch warping; reference selection, optional normal refinement, exposure estimation and on-demand raycasting address distinct alignment failure modes. e-patchese-warpe-visibilitye-visual

Figure 2. One map supplies geometric and photometric constraints to consecutive state updates. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start at the IMU input and follow forward propagation into the LiDAR branch. Scan recombination groups points at image timestamps, and backward propagation compensates motion before point-to-plane residual construction. The blue arrow from LiDAR Update into Visual Update is the key ordering: the image optimizer receives the corrected state and covariance. At the bottom, follow the map retrieval path through visible-voxel queries, raycasting, outlier rejection and reference-patch warping into photometric error construction. The right-hand map stores both LiDAR points and patches. Geometry is appended after the LiDAR update; visual observations and reference patches are updated after image alignment. e-overviewe-filtere-mape-visibilitye-visuale-confige-state

What it supports. The diagram makes the map’s dual role explicit: planes constrain LiDAR registration and support patch geometry, while stored appearance supplies visual constraints. This explains how the system can share depth information without reconstructing an independent visual map. It is an online estimator with map maintenance, rather than a learned predictor of future actions.

Where the evidence stops. The dashed normal-refinement block is optional and disabled in default experiments. The diagram describes information flow, not guaranteed observability when both sensors lose useful constraints; timing and rigid extrinsic calibration are assumed.

2. Motivation

2.1 The problem and the proposed response

Source description

LiDAR loses geometric constraints near single planes, while cameras suffer from weak texture, blur and changing illumination. FAST-LIVO2 seeks complementary constraints without expensive feature matching or separate visual depth reconstruction, under the computational limits of onboard robotics. e-probleme-overview

2.2 What this reading follows

A wall can provide abundant LiDAR returns while still leaving motion along the wall poorly constrained. Image texture can help, but only if geometric projection, brightness and visibility are handled well enough for direct alignment. FAST-LIVO2 organizes these dependencies around a shared voxel map: inertial propagation supplies a prior, LiDAR refines it, and image patches refine it again. This reading follows that information flow, then separates benchmark accuracy from the narrower evidence for patch selection, normal refinement and raycasting. It covers the supplied August 2024 arXiv v2 PDF and its embedded supplement; the catalog’s later journal edition was not available for comparison. e-identitye-probleme-overviewe-filtere-benchmark

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryFoundational work
ArchitectureNot applicable
Prediction paradigmNot applicable
QuadrantNot applicable

This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The recorded foundational state-estimation category is appropriate. A shared voxel map and filter are not a learned One Model world/action architecture. FAST-LIVO2 estimates current state and geometry; the separate planner/controller produces actions, so architecture, prediction paradigm and quadrant remain Not applicable. e-overviewe-filtere-uav

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Timestamped raw LiDAR points, camera images and IMU angular-velocity/acceleration measurements; known sensor timing and pre-calibrated rigid extrinsics.
  • Estimated IMU pose, velocity, biases, gravity and relative inverse camera exposure, with covariance; a local voxel map and registered colored point cloud.

4.2 Equations and their role

p(xyl,yc)p(ycx)p(xyl),p(xyl)p(ylx)p(x)p(\mathbf{x}\mid\mathbf{y}_l,\mathbf{y}_c)\propto p(\mathbf{y}_c\mid\mathbf{x})p(\mathbf{x}\mid\mathbf{y}_l),\qquad p(\mathbf{x}\mid\mathbf{y}_l)\propto p(\mathbf{y}_l\mid\mathbf{x})p(\mathbf{x})
Equations (6)–(7): x is the system state, y_l and y_c are LiDAR and camera measurements, and p(x) is the IMU-propagated prior. The factorization assumes conditional independence of measurement noise. This Bayesian identity does not guarantee identical results after different nonlinear linearizations. e-filtere-sequential
0=τk(Ik(ui+Δu)δIk)τr(Ir(ui+AirΔu)δIr)0=\tau_k\left(I_k(\mathbf{u}_i+\Delta\mathbf{u})-\delta I_k\right)-\tau_r\left(I_r(\mathbf{u}'_i+\mathbf{A}_i^r\Delta\mathbf{u})-\delta I_r\right)
Equation (22): I_k and I_r are measured current/reference intensities, delta I denotes their noise, tau denotes relative inverse exposure, u_i and u_i prime are projected patch centers, and delta u is a current-patch pixel offset. A_i^r warps that offset into the reference patch. Fixing tau_0 = 1 removes the all-zero exposure degeneracy. e-visual

5. Method in detail

5.1 Why a LiDAR update comes before the image update

Source description

At an image timestamp, FAST-LIVO2 first needs a plausible estimate of where its sensors are. IMU integration supplies that estimate, and motion compensation expresses the scan consistently at its endpoint. LiDAR then compares observed points with local planes and updates the state and covariance. The camera starts from that corrected distribution. This matters because direct image alignment searches locally through image gradients: a poor initial projection can lead it toward a wrong alignment. The Bayesian factorization in Equations (5)–(7) explains how independent measurement likelihoods can be applied sequentially. It does not make the resulting nonlinear algorithms interchangeable. The supplement’s AMvalley03 comparison illustrates the practical distinction, but also changes iteration limits and joint residual scaling. Its empirical advantage therefore belongs to the tested implementation and settings. e-statee-filtere-lidare-visuale-sequential

Figure 5. Plane orientation determines how an image patch should warp between views. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read panel (a) from the central reference frame toward the two target frames. The red warping arrows agree with Equation (13): the transform maps reference-patch pixels into each target patch using relative pose and a local plane normal. The visual measurement in Equation (21) uses the reverse, current-to-reference direction; these are distinct uses of warping, not contradictory arrows. In panel (b), the normal on the unit sphere is represented by a scaled vector M satisfying the source-plane constraint. Its two free coordinates m permit unconstrained optimization, after which normalization recovers the plane normal. LiDAR supplies the initial geometric estimate. e-warpe-visuale-affine-studye-confige-ablation

What it supports. Using a plane allows depth to vary within a patch instead of imposing constant depth across its pixels. The supplementary patch projections and map comparisons support this choice in the two examined scenes. Further normal refinement is a separate option; the benchmark’s default system already uses the LiDAR plane prior.

Where the evidence stops. A small planar patch remains an approximation, especially near discontinuities or foliage. The picture illustrates a geometric parameterization, not measured normal accuracy, and the normal-refined variant is not uniformly better in the benchmark.

5.2 How LiDAR geometry becomes a visual constraint

Source description

A visual map point is a LiDAR-derived location with image observations attached to it. Its depth does not have to be triangulated by a separate visual backend. The surrounding plane tells the system how neighboring patch pixels should move between views; the reference patch supplies their appearance. Reference selection balances agreement with other observations against a near-normal viewing direction, instead of automatically choosing the closest view. Optional normal refinement adjusts the plane using photometric consistency across observations and freezes it after convergence. Meanwhile, the main visual residual scales intensities by estimated relative inverse exposure so brightness changes need not be explained entirely as motion. These mechanisms address different sources of mismatch. The reported implementation uses 8×8 patches for alignment and 11×11 for refinement, with refinement disabled by default. e-overviewe-patchese-warpe-visuale-config

5.3 What the evidence says about usefulness for robots

Reader analysis

Reader analysis: the strongest evidence is the combination of trajectory benchmarks, mechanism studies and a working downstream flight stack, each answering a different question. Table II compares odometry errors under the authors’ evaluation choices; it does not measure navigation success. The corridor image shows where extra visual anchors come from, but lacks a quantitative disabled-module control. Normal-change curves explain optimizer behavior without revealing true normal error. In the flight application, FAST-LIVO2 supplies localization and registered points to separate planning and control components, with only Basement and Woods using autonomous planning. Its shared map is therefore useful infrastructure for robot control, but is not itself an action-generating model. A reproduction should keep pose accuracy, map appearance, latency and closed-loop behavior as separate outcomes. e-benchmarke-raycast-studye-normal-studye-uave-flight-modese-runtime

5.4 Training and inference

During training

Reader analysis

The odometry estimator has no learned weights or offline training stage: it performs online filtering and local map optimization. The downstream 3D Gaussian Splatting example has its own training stage and must not be treated as training FAST-LIVO2. e-filtere-warpe-rendering

During inference

Source description

Default experiments enable exposure estimation and reference updates but disable normal refinement. FAST-LIVO2 supplies localization and point clouds to the UAV stack; Bubble planner and MPC produce motion commands. Basement and Woods use autonomous planning, whereas Narrow Opening and SYSU Campus disable planning and use manual flight commands. e-confige-uave-flight-modes

5.5 Implementation flow

  1. Synchronize and predict

    Recombine LiDAR points at camera timestamps. Forward IMU propagation predicts the 19-dimensional state and covariance; backward propagation compensates each scan for motion. The state contains attitude, position, velocity, gyroscope and accelerometer biases, gravity, and inverse exposure relative to the first frame. e-statee-filter

  2. Register geometry

    Transform undistorted points into the map and form point-to-plane residuals in occupied planar voxels. The LiDAR likelihood accounts for point and plane uncertainty, including increased ranging uncertainty from beam divergence at oblique incidence. Iterate the filter, then update map geometry. e-lidare-filter

  3. Maintain shared anchors

    A hash table indexes 0.5 m root voxels with adaptive octrees. Planar leaves hold geometry, uncertainty and LiDAR points; selected points also store patch pyramids. A sliding local map reuses memory. High-gradient candidates populate image grids; reference patches are scored by cross-correlation with other observations and near-normal viewing direction. e-mape-patches

  4. Recover visible map points

    Query voxels hit by the current LiDAR scan and previously visible map points. For unoccupied image cells, sample rays through the existing map until suitable projected points are found. Reject occlusions, depth discontinuities and excessive viewing angles before alignment. e-visibility

  5. Align and refine appearance

    Use plane geometry to warp patches and minimize exposure-scaled photometric discrepancies. An inverse compositional formulation reuses pose Jacobians. The visual update proceeds from coarse to fine, followed by visual-map and reference-patch updates. Optional normal refinement minimizes inter-patch photometric error in a separate thread using a two-variable reparameterization; converged normals and references are fixed. e-warpe-visuale-patches

6. Experiments & results

FAST-LIVO2 estimates sensor motion and builds a colored local map by combining IMU propagation, direct LiDAR registration and sparse image-patch alignment. Its central design is a shared voxel map and a LiDAR-first, camera-second filter update. The reported gains concern odometry and mapping; autonomous flight additionally uses a separate planner and controller.

6.1 Read the original evidence

Table II. Separate the default result from the normal-refined variant before comparing accuracy. Original paper, p. 13 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Rows are individual sequences, grouped by dataset; columns mix external baselines, the authors’ LiDAR-only subsystem and FAST-LIVO2 variants. Lower translational RMSE is better. Read the four rightmost columns carefully: removing exposure estimation, enabling normal refinement, removing reference updating, and the default system are different configurations. The bottom row contains the paper’s reported averages, while the retained cross-symbol footnote marks total failure. The evaluation uses selected single cameras and excludes Hilti’23 Site 3. Baseline implementations also differ from untouched releases: R3LIVE is adapted, SDV-LOAM is reimplemented with a LiDAR backend, and LVI-SAM’s loop closure is disabled. e-benchmarke-ablatione-datasetse-config

What it supports. The reported default average is 0.045 m, compared with 0.137 m for FAST-LIVO, 0.151 m for FAST-LIO2 and 0.278 m for R3LIVE. Normal refinement reports 0.044 m. Removing exposure estimation or reference updating reports 0.051 m or 0.089 m, respectively, making reference updating the larger of these measured ablation effects.

Where the evidence stops. The nearby prose highlights 0.044 m, which belongs to the normal-refined column. Treat the Average row as reported: repeat counts, uncertainty intervals and a detailed aggregation prescription are not supplied. Inspect failures and individual sequences before ranking systems.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Public-dataset odometry

25 NTU-VIRAL, Hilti’22 and Hilti’23 sequences; selected single cameras at 10 Hz; Hilti’23 Site 3 excluded. LVI-SAM loop closure disabled.

Default: 0.045; with normal refinement: 0.044.

Absolute translational RMSE (m), reported Average row

FAST-LIVO: 0.137; FAST-LIO2: 0.151; R3LIVE: 0.278.

Table II separates the two FAST-LIVO2 configurations; adjacent prose highlights 0.044. These reported averages have no repeat-run uncertainty. Baselines include adapted implementations, and failures are marked separately. e-datasetse-confige-benchmark

Photometric-module ablations

Same 25-sequence Table II evaluation.

No exposure estimation: 0.051; no reference update: 0.089; normal refinement enabled: 0.044.

Reported average translational RMSE (m)

Default: 0.045.

Reference updating has the largest reported ablation effect. Normal refinement changes the mean by only 1 mm and can worsen individual sequences with dim or blurred images. e-benchmarke-ablation

ESIKF update strategy on AMvalley03

MARS-LVIG aerial sequence with RTK ground truth; desktop platform; strategy-specific iteration limits and tuned joint-update scaling.

Synchronous sequential: 0.68 m; 23.1 ms.

APE RMSE (m); mean processing time (ms)

Asynchronous standard: 3.12 m; 27.6 ms. Synchronous standard: 2.45 m; 49.9 ms.

Supports this implementation on one sequence. Recombination, update ordering, iteration limits and scaling prevent attribution to ordering alone. e-sequentiale-config

Odometry processing time

Table III public and private sequences; Intel i7-10700K/32 GB RAM and RB5 Kryo585/8 GB RAM.

Desktop: 30.03, comprising 17.13 LiDAR and 12.90 image; ARM: 78.44.

Reported average processing time per LiDAR/image pair (ms)

FAST-LIVO: 41.43; R3LIVE: 108.36. LiDAR-only FAST-LIO2 reports 19.68 in Table S3.

Average latency supports 10 Hz operation on the tested platforms. Failure-dependent sequence coverage and different sensor workloads limit aggregate speed rankings; means do not establish worst-case deadlines. e-confige-runtimee-lio-runtime

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure S3. Reference quality and normal refinement are related but separately testable choices. Original paper, p. 22 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Begin with regions A–D on HIT Graffiti Wall and E–H on HKU Centennial Garden. The numbered camera poses identify the observations shown in each patch row; red boxes mark selected reference patches. Compare how the same surface’s apparent resolution and orientation change across views. These displayed patches are 40×40 for visualization, whereas the normal-refinement experiment uses 11×11 patches. Next read panel (c): its horizontal axis is iteration number and its vertical axis is angular change from the initial LiDAR-derived normal. The A–H legend connects each curve to its scene region. It does not plot error against a known true normal. e-normal-studye-patchese-warpe-benchmarke-ablation

What it supports. Structured regions generally need smaller angular corrections than foliage-related regions in these selected examples. The reference-selection mechanism favors representative appearance and a more frontal view, while refinement adjusts local geometry from multiple observations. Together the panels help explain the design; the benchmark table is needed to assess whether those choices improve trajectory accuracy.

Where the evidence stops. The curves continue to fluctuate after their initial rise despite the authors’ convergence language. Neither stabilization nor a larger correction establishes a more accurate normal. These eight selected regions lack ground-truth normal errors and do not establish universal benefit.

Figure S4. Raycasting recalls existing visual anchors when recent LiDAR returns are sparse. Original paper, p. 23 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Follow the cyan trajectory through the corridor and the yellow arrows indicating travel direction. The dashed box links the turn toward a wall to the camera inset. The original caption’s color key is essential: blue/cyan dots are visual map points recalled by raycasting, yellow dots come from voxel queries, and red denotes current LiDAR scan points. At close range the recent scans contain few points, so their voxel hits alone retrieve few visual anchors. Section VII-A therefore casts rays only through image cells still lacking selected map points. Those rays search stored geometry; they do not measure new depth or create observations of unseen surfaces. e-raycast-studye-visibility

What it supports. The inset demonstrates that many stored visual anchors can remain available even when LiDAR returns become scarce. This is a concrete explanation for why map retrieval must not rely entirely on the latest scan. It supports the mechanism in this corridor example, while leaving the magnitude of its trajectory-level benefit unmeasured.

Where the evidence stops. There is no quantitative raycasting-on versus raycasting-off result here. Recalled points require previously mapped surfaces and useful image gradients; their presence alone does not prove correct correspondences or full pose observability.

Figure S6. Update strategy changes the linearization conditions in this aerial mapping comparison. Original paper, p. 24 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Compare corresponding columns: (a) is asynchronous standard updating, (b) synchronizes scans and images but updates their measurements jointly, and (c) synchronizes them and applies LiDAR before vision. The blue and orange boxes select corresponding road regions, enlarged along the bottom for inspection of alignment and layering. The cyan line is the flight path; red points in (c3) mark the current LiDAR scan near the challenging slope. Read these images alongside the same page’s RTK-based APE and runtime results. The numerical comparison belongs to this single MARS-LVIG sequence and to the iteration budgets and measurement scaling described above the figure. e-sequentiale-filtere-config

What it supports. The page reports APE RMSE of 3.12 m, 2.45 m and 0.68 m for (a), (b) and (c), respectively; corresponding times are 27.6, 49.9 and 23.1 ms. The authors attribute the sequential variant’s advantage to the LiDAR-corrected visual prior and avoiding repeated LiDAR fusion at every visual pyramid level.

Where the evidence stops. The joint update uses fewer iterations per pyramid level and a tuned normalization factor. Thus this is not a controlled test of ordering alone. Conditional-independent Bayesian equivalence does not remove practical differences from nonlinear linearization, synchronization or computational budgets.

7. Analysis & limitations

7.1 What the evidence leaves open

Source description

Long-distance drift remains because loop closure and sliding-window optimization are future work. Accurate timing/extrinsics are assumed. Direct alignment still needs a useful prior; degraded images can reduce accuracy, and exposure modeling omits camera response and vignetting. e-conclusione-statee-probleme-ablatione-exposure

Reader analysis

Private-sequence return-to-origin checks and visual map sharpness do not replace trajectory-wide ground truth. The raycasting example is qualitative; normal-angle stabilization is not ground-truth normal accuracy. Reported benchmark tables omit repeat counts and uncertainty intervals. e-privatee-raycast-studye-normal-studye-benchmark

7.2 Questions for discussion

  1. Would sequential fusion retain its advantage with matched iteration budgets and calibrated measurement weighting?
  2. Can confidence in texture, exposure and normal estimates prevent visual updates from worsening a strong LiDAR estimate?

8. Reproducibility audit

8.1 Requirements and known gaps

Source description

Reproduction needs calibrated synchronized streams, matching camera models and sensor noise settings. Section IX-A specifies C++/ROS, 1:3 temporal LiDAR subsampling, octree depth 3 and photometric noise 100; it distinguishes 8×8 alignment patches from 11×11 refinement patches. Exact software versions and code revision are unspecified. e-statee-config

Reader analysis

Some plane criteria and convergence details are delegated to prior work. The visual text specifies coarse-to-fine processing, but Algorithm 1 counts levels upward from zero while Section V-D calls zero the finest level; verify the indexing when implementing. Public benchmark results also depend on documented R3LIVE adaptations and the SDV-LOAM reimplementation. e-mape-filtere-patchese-visuale-benchmark

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Isolate sequential fusion from its computation budget

Reader-proposed check, not performed: replay AMvalley03 with the paper’s asynchronous-standard, synchronous-standard and synchronous-sequential settings first. Then compare the two synchronous variants with identical input points, initial state/covariance, sensor noise, outlier gates and convergence tolerances, reporting both matched iteration limits and matched elapsed-time budgets. Keep and disclose the joint residual scaling, including a sweep around the reported 0.0032. Measure RTK-based APE, per-frame latency, actual iteration counts and visual residuals before and after updating. A persistent sequential advantage under these controls would support the corrected-prior explanation; disappearance of the gap would implicate the original budget or weighting choices. e-sequentiale-filtere-config

Check 2: Test whether recalled anchors prevent tracking loss

Reader-proposed check, not performed: replay Narrow Corridor with raycasting enabled and disabled while holding voxel-query logic, map initialization, patches, exposure estimation and all other settings fixed. Compare the wall-facing turn with earlier segments where LiDAR returns are plentiful. Log scan-point counts, visual anchors retrieved by each route, accepted residuals, tracking-loss duration and latency. Report endpoint drift separately; add trajectory APE only if independent synchronized ground truth is obtained. The mechanism predicts an interaction: raycasting should help most during sparse-return intervals. More recalled points without better tracking, or an equal benefit in ordinary segments, would weaken that explanation. e-raycast-studye-visibilitye-privatee-config

8.3 Reading coverage

Visual audit: All 30 pages of the supplied PDF were rendered and visually inspected, including the title/author/version block, complete method and equations, Tables I–III, Figures 1–16, and embedded supplementary Figures S1–S15 and Tables S1–S3. All six final original crops were inspected after cropping; the narrow Figure 5 crop uses a fresh 300-DPI render. Figure 2 flow and Figure 5 warp directions were cross-checked against the equations and algorithm. Benchmark configuration distinctions and the pyramid-indexing ambiguity are preserved in the report. All supporting pages for the retained method, numerical, evaluation and reproduction claims are declared here. External videos, code, datasets and separately linked supplements remain outside this reading.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Abstract; I. Introduction; II. Related Works
  • III. System Overview; IV. Error-State Iterated Kalman Filter with Sequential State Update, A–D
  • V. Local Mapping, A–E; VI. LiDAR Measurement Model, A–B; VII. Visual Measurement Model, A–B
  • VIII. Datasets for Evaluation; IX. Experiment Results, A–E
  • X. Applications, A–C; XI. Conclusion and Future Work; References
  • Embedded Supplementary Material: I. System Module Validation, A–E; II. Additional Information; all supplementary figures and tables

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • Identity/edition note: the inspected title page identifies arXiv:2408.14035v2 [cs.RO], 28 August 2024. Its title and all 14 authors match the catalog. The catalog describes a 2025 IEEE Transactions on Robotics publication; that journal edition was not supplied, and equivalence between editions is unverified.
  • Acquisition omission retained: text extraction does not reconstruct figure images; the retained PDF was therefore visually inspected on all 30 pages.
  • Acquisition omission retained: separate supplemental material availability has not been fully verified. The ten supplementary pages embedded in this PDF were read and visually inspected; separately linked supplements and videos were not inspected.
  • Code, datasets and referenced external implementations were not inspected; no experiments were reproduced.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title, author block, affiliations and arXiv marginInspect

Title and 14 authors match the supplied identity. The page labels the artifact arXiv:2408.14035v2, 28 August 2024, and lists University of Hong Kong MaRS and Information Science Academy of China Electronics Technology Group Corporation.

Go to primary source ↓
e-problemPDF pp. 2–3, Sections I and II-AInspect

Describes complementary sensor degeneration, computation constraints and the reliance of direct methods on accurate priors.

Go to primary source ↓
e-overviewPDF pp. 4–5, Section III and Figure 2Inspect

Shows IMU propagation, LiDAR then visual updates, shared voxel-map retrieval/update paths and separate-thread normal refinement.

Go to primary source ↓
e-statePDF p. 5, Section IV-A, Table I and Equations (1)–(2)Inspect

Defines calibrated timing/extrinsics, the 19-dimensional inertial state, relative inverse exposure, and random-walk noise models.

Go to primary source ↓
e-filterPDF pp. 6–7, Sections IV-B–D, Algorithm 1 and Equations (3)–(11)Inspect

Defines scan recombination, forward/backward propagation, conditional-independent Bayesian fusion and iterated state/covariance updates. Algorithm 1 increments level from zero.

Go to primary source ↓
e-mapPDF pp. 7–8, Sections V-A–B and Figure 4Inspect

Defines root voxels, planar octree leaves, memory reuse and geometry updates; delegates detailed plane criteria and convergence to reference [14].

Go to primary source ↓
e-patchesPDF p. 8, Sections V-C–D and Equation (12)Inspect

Defines image-grid candidate selection, patch observations and reference scores combining NCC with viewing direction; identifies pyramid level zero as highest resolution.

Go to primary source ↓
e-warpPDF pp. 8–9, Section V-E, Figure 5 and Equations (13)–(16)Inspect

Defines reference-to-target warping, photometric normal optimization, two-dimensional reparameterization and freezing of converged normals/reference patches.

Go to primary source ↓
e-lidarPDF pp. 9–10, Section VI, Equations (17)–(20) and Figure 6Inspect

Defines point-to-plane measurements with point/plane uncertainty and beam-divergence-dependent range uncertainty.

Go to primary source ↓
e-visibilityPDF pp. 10–11, Section VII-A and Figures 7–8Inspect

Describes voxel queries, raycasting into unoccupied image cells, first suitable voxel termination, and occlusion/depth/view-angle rejection.

Go to primary source ↓
e-visualPDF p. 11, Section VII-B, Equations (21)–(23)Inspect

Defines current-to-reference warp in the photometric residual, inverse compositional pose updates, relative exposure anchoring and coarse-to-fine alignment.

Go to primary source ↓
e-datasetsPDF pp. 11–12, Section VIII-AInspect

Describes camera selection and rates, ground-truth sources, Hilti evaluation service and exclusion of four Hilti’23 Site 3 sequences.

Go to primary source ↓
e-privatePDF p. 12, Section VIII-B; PDF p. 25 (supplement p. 5), Table S1 and footnotesInspect

Describes 20 private sequences and return-to-origin collection. Table S1 reports 66 min 50 sec; body text summarizes duration as 66.9 min.

Go to primary source ↓
e-configPDF pp. 12–13, Section IX-A and default-configuration paragraph in IX-BInspect

Specifies default exposure/reference updates with normal refinement off, patch sizes, subsampling, voxel/noise settings, C++/ROS and desktop/ARM hardware.

Go to primary source ↓
e-benchmarkPDF p. 13, Table II, Average row and Section IX-BInspect

Reports default/normal-refined averages 0.045/0.044 m, ablations 0.051/0.089 m, and baseline averages. Describes loop-closure removal and baseline adaptations. No repeat-run uncertainty is supplied.

Go to primary source ↓
e-ablationPDF pp. 13–14, Table II per-sequence entries and Section IX-B discussionInspect

Exposure/reference removal worsens reported means by 6/44 mm; normal refinement improves the mean by 1 mm but sometimes degrades dim/blurred sequences.

Go to primary source ↓
e-runtimePDF p. 15, Section IX-E; PDF p. 17, Table III, Average row and failure footnoteInspect

Reports 30.03 ms desktop, 78.44 ms ARM, 41.43 ms FAST-LIVO and 108.36 ms R3LIVE; failures are separately marked.

Go to primary source ↓
e-lio-runtimePDF p. 25 (supplement p. 5), Table S3, Overall Average and failure footnoteInspect

Reports 19.68 ms for LiDAR-only FAST-LIO2; numerous sequence entries are failures.

Go to primary source ↓
e-uavPDF pp. 17–18, Section X-A and Figure 12Inspect

Separates FAST-LIVO2 localization/map outputs from Bubble planner, MPC and low-level flight control; distinguishes autonomous and manual demonstrations.

Go to primary source ↓
e-flight-modesPDF p. 25 (supplement p. 5), Table S2 and footnote 6Inspect

Basement/Woods are autonomous with planning; Narrow Opening/SYSU Campus are manual with FAST-LIVO2 and MPC but no planning.

Go to primary source ↓
e-renderingPDF pp. 18–19, Section X-C and Figures 15–16Inspect

Uses estimated poses/maps for meshing, texturing and separate 3DGS training; the random-frame rendering comparison is a downstream application.

Go to primary source ↓
e-conclusionPDF p. 19, Section XIInspect

Acknowledges long-distance odometry drift and proposes loop closure, sliding-window optimization and semantic mapping as future work.

Go to primary source ↓
e-affine-studyPDF pp. 21–22 (supplement pp. 1–2), Section I-A and Figures S1–S2Inspect

Compares constant-depth warping, LiDAR plane priors and refined normals on CBD Building 02 and Office Building Wall; displayed map/patch comparisons favor plane-aware warping.

Go to primary source ↓
e-normal-studyPDF pp. 21–22 (supplement pp. 1–2), Section I-B and Figure S3Inspect

Regions A–H link five observations to selected reference patches and angular changes from initial normals. Visualization patches are 40×40; normal refinement uses 11×11 patches.

Go to primary source ↓
e-raycast-studyPDF p. 23 (supplement p. 3), Section I-C and Figure S4/captionInspect

Shows a 1.9 m Narrow Corridor: blue points come from raycasting, yellow from voxel queries and red from the LiDAR scan. No quantitative disabled-raycasting comparison is supplied.

Go to primary source ↓
e-exposurePDF p. 23 (supplement p. 3), Section I-D and Figure S5; PDF p. 30 (supplement p. 10), Figure S15Inspect

Exposure diagnostics compare synthetic scaling and camera-API readings; authors attribute occasional mismatches to unmodeled response/vignetting. UAV plots also show exposure estimates.

Go to primary source ↓
e-sequentialPDF p. 24 (supplement p. 4), Section I-E and Figure S6/captionInspect

AMvalley03 strategies report 3.12/2.45/0.68 m APE and 27.6/49.9/23.1 ms. Joint updates use at most 3 iterations per pyramid level versus up to 5 otherwise, with joint scale 0.0032.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.