PAPER REPORTENAll readings ↗

LIC-Fusion: LiDAR-Inertial-Camera Odometry

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Xingxing Zuo; Patrick Geneva; Woosik Lee; Yong Liu; Guoquan Huang

Affiliations: Institute of Cyber-System and Control, Zhejiang University, Hangzhou, China; Department of Computer and Information Sciences, University of Delaware, Newark, DE 19716, USA; Department of Mechanical Engineering, University of Delaware, Newark, DE 19716, USA

Source: IROS 2019 · ref-c7de129d1b2c440e37ac ↗ · Catalog record

Reading: 524 / 558 · 5 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: LIC-Fusion combines sparse visual and LiDAR constraints with inertial propagation and online calibration, improving reported odometry accuracy while leaving the contribution of each component unisolated. identitymotivationcloninglidar-geometryoutdoor-ateindoor-resultsreporting-gaps

At a glanceWhat to know
Research problem
Source description

Cameras lose useful information under poor lighting, LiDAR supplies lighting-independent ranges but sparse, slower observations, and unaided IMUs drift. The problem is robust six-degree-of-freedom motion estimation from these complementary, asynchronous sensors. The authors seek to retain measurement correlations through tight fusion while adapting spatial and temporal calibration. motivationstate-calibration

Core mechanism
Author claim

The authors propose a single-thread MSCKF extension that jointly estimates motion and calibration using sparse visual observations and two LiDAR feature types. They describe fusion as optimal up to linearization errors; the experiments do not independently establish this optimality claim. motivationcompression

A key reported resultOutdoor trajectory estimation: Measured LIC-Fusion average ATE: 4.06 m; one sigma: 3.42 m.

Reported average ATE and one-sigma variability, in meters; lower ATE is better.. One approximately 800 m, four-minute robot sequence; RTK GPS reference; six runs per algorithm; trajectories aligned by a best-fit transform minimizing overall error.

MSCKF: 10.75 m, sigma 3.56 m; LOAM: 23.08 m, sigma 2.63 m. Fusion has the smallest reported average ATE on this sequence. The sigma row is variability, not a confidence interval. LOAM uses its global-map output, unlike sliding-window LIC-Fusion. outdoor-ateexperiment-setup

Reading caution
Reader analysis

Figure 4 is called mean squared error in its caption and prose, but its axis reads Error (m). The prose claims reductions of 2.5 m and 5 m, which do not match subtraction of Table I ATEs. These definitions remain unresolved; retain Table I values without treating the plot as a verified MSE benchmark. outdoor-ambiguity

Core contributions

  • Author claim

    The authors propose a single-thread MSCKF extension that jointly estimates motion and calibration using sparse visual observations and two LiDAR feature types. They describe fusion as optimal up to linearization errors; the experiments do not independently establish this optimality claim. motivationcompression

  • Source description

    LiDAR distance constraints include pose and extrinsic-calibration dependence, while timestamp dependence enters through cloned IMU states. This provides an explicit route for geometric measurements to refine calibration. cloninglidar-geometrylidar-noise

Figure 1. Sparse scene geometry supplies two kinds of LiDAR constraint alongside visual tracks. Original paper, p. 1 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with the legend: red points identify LiDAR edges and blue points identify planar surf features. The green line is the estimated trajectory, as specified by the original caption; the inset shows the camera view with sparse tracked features. Connect this picture to Section II-D rather than reading it as an architecture diagram. High-curvature scan regions supply edges, and low-curvature regions supply planes. After projection into the previous scan, an edge match defines a point-to-line distance and a surf match defines a point-to-plane distance. Camera tracks follow their own reprojection measurement model before entering the same estimator. motivationlidar-geometrylidar-noisevisual-updatecompression

What it supports. The illustration makes the sparsity choice concrete: the estimator uses selected geometric features from scans and tracked image features to constrain motion. Their measurement models differ, but both contribute to the shared state update. The green trajectory represents estimated motion; the figure does not show a planned route or action output.

Where the evidence stops. The display contains no data-flow arrows or filter blocks. Feature colors agree with the caption and Section II-D. It does not establish matching accuracy, calibration convergence or the separate benefit of edges versus planes.

2. Motivation

2.1 The problem and the proposed response

Source description

Cameras lose useful information under poor lighting, LiDAR supplies lighting-independent ranges but sparse, slower observations, and unaided IMUs drift. The problem is robust six-degree-of-freedom motion estimation from these complementary, asynchronous sensors. The authors seek to retain measurement correlations through tight fusion while adapting spatial and temporal calibration. motivationstate-calibration

2.2 What this reading follows

LIC-Fusion asks how a camera, LiDAR and IMU can jointly track motion when their strengths and failure modes differ. Its answer is a geometric filter: inertial measurements advance the state, sparse image and scan features constrain that state, and sensor transforms and clock offsets are refined alongside motion. Read the feature illustration first, then compare the outdoor ATE table with the error curve, and finally examine the indoor endpoint results beside the raw motion diagnostic. The key distinction is between evidence that the combined estimator works on these sequences and evidence identifying why it works. This six-page arXiv v2 provides the former, with several reporting boundaries. identitymotivationcloninglidar-geometryoutdoor-ateindoor-resultsreporting-gaps

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryFoundational work
ArchitectureNot applicable
Prediction paradigmNot applicable
QuadrantNot applicable

This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The recorded foundational-work/state-estimation category fits the explicit physical state, geometric residuals and Kalman updates. Architecture, prediction paradigm and quadrant are correctly not applicable to a world-action-model taxonomy: the method contains neither a learned future/action predictor nor inverse-dynamics action extraction. Combining three sensors in one estimator does not establish a One Model architecture. state-calibrationlidar-geometryvisual-updatecompression

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • IMU angular-rate and linear-acceleration measurements
  • Monocular images with FAST features tracked using KLT
  • LiDAR scans reduced to high-curvature edge and low-curvature planar surf features
  • Estimated IMU orientation, position, velocity and biases, with covariance
  • Camera–IMU and LiDAR–IMU rigid transforms and time offsets
  • Sliding windows of pose clones at image and LiDAR observation times

4.2 Equations and their role

tI=tC+tdC,tI=tL+tdLt_I=t_C+t_{dC},\qquad t_I=t_L+t_{dL}
Equations (7)–(8): t_I is IMU-clock time; t_C and t_L are reported camera and LiDAR times; t_dC and t_dL are their estimated offsets. The IMU clock is the reference. state-calibration
r(Ll+1pfi)=(LlpfiLlpfj)×(LlpfiLlpfk)2LlpfjLlpfk2r({}^{L_{l+1}}\mathbf p_{f_i})=\frac{\left\|({}^{L_l}\mathbf p_{f_i}-{}^{L_l}\mathbf p_{f_j})\times({}^{L_l}\mathbf p_{f_i}-{}^{L_l}\mathbf p_{f_k})\right\|_2}{\left\|{}^{L_l}\mathbf p_{f_j}-{}^{L_l}\mathbf p_{f_k}\right\|_2}
Equation (24): the current feature i is projected into the previous LiDAR frame L_l; old features j and k define its corresponding line. The Euclidean norm of the cross product divided by line-segment length gives the point-to-line distance. lidar-geometry
rc=THδx+nc,nc=QH1n\mathbf r_c=\mathbf T_H\delta\mathbf x+\mathbf n_c,\qquad \mathbf n_c=\mathbf Q_{H1}^{\top}\mathbf n
Equation (33): r_c is the compressed residual, T_H the retained factor from QR of the stacked measurement Jacobian, delta x the state error, and n_c the projected measurement noise. Q_H1 is the retained QR basis and n the stacked noise. This reduces update size without introducing a learned representation. compression

5. Method in detail

5.1 Put asynchronous measurements on the IMU clock

Source description

Begin with the state, not the sensors as separate odometry systems. LIC-Fusion stores IMU orientation, position, velocity and biases, plus camera and LiDAR transforms and time offsets. It also stores two windows of historical pose clones, one associated with images and one with scans. When a scan arrives, its reported time is shifted by the current offset estimate, and the IMU state is propagated to that corrected time before cloning. Equation (19) makes the new clone sensitive to timing error through angular velocity and global linear velocity. This dependence is carried into the augmented covariance. Later geometric measurements can therefore influence timing through their dependence on those clones. Camera arrivals use the analogous construction; offline extrinsic initialization is followed by online refinement. state-calibrationcloningexperiment-setup

5.2 Convert sparse structure into constraints on the shared state

Source description

LiDAR processing begins by selecting high-curvature edges and low-curvature surf features. The current scan is projected into the previous frame using the estimated motion and extrinsic transform. KD-tree lookup provides candidate matches. Two old edge points, with the second on a neighboring scan ring, define a line; three surf points define a plane. Their distance residuals constrain cloned poses and LiDAR extrinsics. Because these distances are derived from noisy points, the filter propagates point covariance into distance uncertainty before applying its Mahalanobis gate. Images take a different route: FAST features are tracked by KLT, triangulated when lost or spanning the window, and converted to reprojection constraints. Nullspace projection removes explicit dependence on the triangulated feature position, leaving constraints on the estimator state. lidar-geometrylidar-noisevisual-update

5.3 Compress the update, then separate accuracy from mechanism

Reader analysis

After feature processing, the paper stacks the available residuals and Jacobians, assumes independent measurement noise, and applies thin QR using Givens rotations. The retained factor and projected noise define the smaller measurement system used for the standard EKF state and covariance update. This is an estimation pipeline, with no learned representation or action output. Reader interpretation: the reported accuracy comparisons test the resulting system as a whole, not the necessity of each step. Table I favors LIC-Fusion outdoors, whereas Table II favors different methods on different indoor rows. Figure 6 documents the motion challenge but does not remove a component. A causal claim about calibration, LiDAR geometry or compression would require a matched intervention that the supplied experiments do not provide. compressionexperiment-setupoutdoor-ateindoor-resultsmotion-diagnosticreporting-gaps

5.4 Training and inference

During training

Reader analysis

There is no learned model, training loss, pretrained component or training split in this formulation. Sensor extrinsics are calibrated offline and refined online; this is parameter estimation within a geometric filter. state-calibrationcompressionexperiment-setup

During inference

Reader analysis

Runtime processing alternates inertial propagation with image/scan measurement updates. The system estimates a sliding-window trajectory without a global map or loop closures. Its output is localization; the paper specifies no action decoder, planning rollout or executed control policy. cloningcompressionexperiment-setup

5.5 Implementation flow

  1. Propagate a shared physical state

    The state contains current IMU quantities, both sensor calibration blocks, and separate camera-time and LiDAR-time pose windows. Continuous-time inertial kinematics propagate motion and covariance; gyroscope and accelerometer biases follow random walks. Orientation uses JPL quaternions and a local error-state update. state-calibrationcloning

  2. Clone at corrected observation times

    For a new scan, propagate to its timestamp plus the estimated LiDAR offset, then clone orientation and position. Covariance augmentation includes sensitivity to the offset through angular and linear velocity. Camera arrivals use the analogous procedure. cloning

  3. Turn LiDAR matches into geometric residuals

    Project current features into the previous scan and use KD-tree nearest-neighbor lookup. An edge uses two old points, with the second on an adjacent scan ring, to define a line; a surf feature uses three old points to define a plane. Their point-to-line or point-to-plane discrepancies supply linearized EKF constraints. Correspondences are assumed to share a physical edge or plane. lidar-geometrylidar-noise

  4. Weight, reject and combine observations

    Propagate raw point covariances into distance uncertainty and apply a Mahalanobis chi-squared gate. For visual tracks lost or spanning the window, triangulate and form reprojection constraints; MSCKF nullspace projection removes feature-position dependence. Stack residuals, assume independent measurement noise, compress with Givens-based thin QR, and update state and covariance. lidar-noisevisual-updatecompression

6. Experiments & results

LIC-Fusion estimates motion by combining inertial propagation, sparse camera tracks and LiDAR edge/plane constraints in one MSCKF estimator, while refining sensor transforms and clock offsets online. Its strongest outdoor evidence is a reported average ATE of 4.06 m, versus 10.75 m for MSCKF and 23.08 m for LOAM. Indoor results favor fusion on the harder sequences but favor LOAM on A and B; endpoint-only evaluation limits the conclusion.

Source and visual limitations
Reader analysis

The supplied paper has no architecture block diagram and no controlled component, calibration or compression ablation. Figure 1 is its original method illustration, showing the selected features; Figure 6 supplies a motion-input diagnostic in the ablation section, not an ablation result. No calibration-convergence or runtime plots are supplied. The edition therefore uses five original visuals: one feature illustration, both quantitative tables, an outdoor error plot and the IMU diagnostic. motivationmotion-diagnosticreporting-gaps

6.1 Read the original evidence

Table I. LIC-Fusion has the smallest average ATE on the supplied outdoor route. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read across the Average ATEs row first: MSCKF reports 10.75 m, LIC-Fusion 4.06 m and LOAM 23.08 m. The second row reports one-sigma variability, respectively 3.56 m, 3.42 m and 2.63 m; it is not another accuracy metric to add to the first row. Section III-A describes one approximately 800 m outdoor sequence lasting four minutes, with RTK GPS as reference and six runs per algorithm. The trajectories are aligned with a best-fit transform minimizing overall error. Section III also matters: the compared LOAM output uses a global map, while LIC-Fusion maintains a sliding window. outdoor-ateexperiment-setup

What it supports. The table supports a lower average ATE for LIC-Fusion than both compared implementations on this route. It also prevents a misleading certainty claim: LIC-Fusion has 3.42 m reported variability, and the experiment repeats one sequence rather than sampling many independent routes. Accuracy and variability should be read separately.

Where the evidence stops. These are six-run results from one outdoor route, not a multi-dataset aggregate. The paper labels the second row as one sigma, without providing confidence intervals. Global-map LOAM and sliding-window LIC-Fusion also differ in available information.

Figure 4. The outdoor error curve shows changing method rankings and substantial late LOAM drift. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Follow time from left to right, using the legend to identify blue LIC, dotted red MSCKF and magenta LOAM. The plot complements Table I by showing when differences emerge: the magenta trace is low early and rises strongly late, while the blue trace remains comparatively low late in the route. Do not read the visual as evidence that LIC is best at every instant. The horizontal axis is time in seconds, and the vertical axis explicitly says Error (m). Compare that label with the original Figure 4 caption and Section III-A, both of which call the plotted quantity mean squared error. outdoor-ambiguityoutdoor-ate

What it supports. The visual supports the qualitative observation that LIC-Fusion limits late error growth on this route relative to the shown baselines. It also shows why an aggregate ranking conceals temporal variation. Exact aggregate performance should be taken from Table I, whose ATE row explicitly supplies units and values.

Where the evidence stops. The caption/prose say MSE but the axis says meters, leaving the metric unresolved. The prose reductions of 2.5 m and 5 m also differ from Table I ATE subtraction. No corrected MSE values are inferred here.

Table II. Indoor endpoint errors favor LOAM on A and B, and LIC-Fusion on C and D. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each row as a separate handheld sequence. The parenthetical 39 m, 86 m, 55 m and 189 m values are route lengths, not error units or normalization factors. Section III-B explains the evaluation: without indoor trajectory ground truth, the experiment returns the rig to its initial location and measures start-end error. LOAM gives the smallest entries for A and B. LIC-Fusion gives the smallest entries for C and D; C is the sequence explicitly described as vigorous shaking. The accompanying Figure 5 on this page shows estimated paths but provides no ground-truth path that would validate their entire shapes. indoor-protocolindoor-resultsmotion-diagnostic

What it supports. For Indoor-C, the reported values are 49.94 for MSCKF, 1.55 for LIC-Fusion and 2.44 for LOAM. For D they are 46.03, 3.68 and 5.99. These comparisons support a conditional robustness advantage for fusion, while A and B directly rule out a claim of universal superiority over LOAM.

Where the evidence stops. Table II omits units for its error columns, so the values remain unconverted. Endpoint error cannot certify the intervening trajectory. The source gives no indoor uncertainty or repetition count and does not assign measured light levels to individual rows.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Outdoor trajectory estimation

One approximately 800 m, four-minute robot sequence; RTK GPS reference; six runs per algorithm; trajectories aligned by a best-fit transform minimizing overall error.

Measured LIC-Fusion average ATE: 4.06 m; one sigma: 3.42 m.

Reported average ATE and one-sigma variability, in meters; lower ATE is better.

MSCKF: 10.75 m, sigma 3.56 m; LOAM: 23.08 m, sigma 2.63 m.

Fusion has the smallest reported average ATE on this sequence. The sigma row is variability, not a confidence interval. LOAM uses its global-map output, unlike sliding-window LIC-Fusion. outdoor-ateexperiment-setup

Indoor return-to-start localization

Handheld sequences A/B/C/D of 39/86/55/189 m; varied lighting and motion. No indoor trajectory ground truth; the rig returns to its initial location.

Measured LIC-Fusion values A/B/C/D: 0.98 / 1.04 / 1.55 / 3.68.

Average trajectory start-end error; lower is better. Table II omits an explicit unit for the error columns.

MSCKF: 0.99 / 1.55 / 49.94 / 46.03; LOAM: 0.66 / 0.46 / 2.44 / 5.99.

LOAM leads on A/B; fusion leads on C/D. Indoor-C includes deliberate vigorous shaking. Endpoint agreement does not measure whole-trajectory accuracy, and indoor repetition counts and uncertainty are not specified. indoor-protocolindoor-resultsmotion-diagnostic

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure 6. Raw IMU traces document the motion challenge in Indoor-C. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the two panels as inputs to the estimator. The upper panel plots angular velocity in radians per second; the lower plots linear acceleration in meters per second squared. Their legends separate the x, y and z components, and both horizontal axes are time in seconds. The oscillatory intervals and bursts supply visual context for Section III-B, where the authors say they shook the rig vigorously. Then return to the Indoor-C row of Table II to see the corresponding endpoint errors. The plot itself contains neither an estimated pose nor a comparison between algorithms, so its curves must not be interpreted as error reductions. motion-diagnosticindoor-resultscloningreporting-gaps

What it supports. Figure 6 substantiates the qualitative description of Indoor-C as a demanding motion sequence. Together with Table II, it places the reported localization result in context: the input was visibly dynamic. It does not determine which sensor, calibration variable or filtering choice was responsible for the reported outcome.

Where the evidence stops. This is a diagnostic, not a controlled ablation. Raw IMU signals do not isolate calibration, visual degradation or feature matching. The paper supplies no synchronized tracking-error or feature-dropout trace linking individual bursts to estimator behavior.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Figure 4 is called mean squared error in its caption and prose, but its axis reads Error (m). The prose claims reductions of 2.5 m and 5 m, which do not match subtraction of Table I ATEs. These definitions remain unresolved; retain Table I values without treating the plot as a verified MSE benchmark. outdoor-ambiguity

Reader analysis

The evidence covers one outdoor route and four indoor sequences, with no controlled calibration, sensor-removal or compression ablation. It cannot isolate which component causes the improvement. Indoor endpoint evaluation and absent per-sequence lighting labels also constrain robustness claims. outdoor-ateindoor-protocolindoor-resultsreporting-gaps

Reader analysis

Nearest matches are assumed to represent the same physical geometry, and stacked measurement noises are treated as independent. The reported tests do not isolate failures of these assumptions. The authors leave adding camera/LiDAR loop closures to future work. lidar-geometrylidar-noisecompressionfuture-work

7.2 Questions for discussion

  1. How much of the gain survives when calibration refinement is disabled but all measurement residuals remain?
  2. Would the indoor ranking persist under full-trajectory ground truth rather than endpoint error?
  3. How sensitive is the filter to shared LiDAR correspondences that challenge independent-noise assumptions?

8. Reproducibility audit

8.1 Requirements and known gaps

Source description

Match the documented rig: Xsens MTi-300 AHRS IMU, Velodyne VLP-16 LiDAR and monochrome global-shutter Blackfly BFLY-PGE-23S6M camera. Reproduce offline extrinsic initialization, online refinement, the stated baseline variants and the RTK alignment protocol. experiment-setupoutdoor-ate

Reader analysis

The PDF does not give numerical clone-window sizes, curvature thresholds, point/noise covariance settings, chi-squared gate level, full calibration initialization procedure, sensor rates, processing hardware, software versions or runtime measurements. No dataset-release instructions are supplied. These omissions prevent an exact configuration-level reproduction from the paper alone. reporting-gaps

Reader analysis

Proposed checks: compare fixed versus online calibration under controlled timestamp/transform perturbations, and remove LiDAR edge or plane residuals under matched replay conditions. Measure calibrated parameter error and trajectory error, including full-path ground truth where possible; neither check was performed for this report. state-calibrationcloninglidar-geometryindoor-protocol

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Perturb timestamps and test whether calibration recovers

Reader-proposed experiment, not a reported reproduction: replay identical camera/LiDAR/IMU streams with independently injected camera or LiDAR offsets of 0, ±10 and ±30 ms. Compare a correctly fixed offset control, an incorrectly fixed offset, and online offset estimation initialized at the same incorrect value. Keep feature selection, noise settings, clone windows and random seeds matched. Measure recovered offset error, convergence time and RTK-aligned trajectory ATE under one common evaluation protocol. Repeat with controlled extrinsic perturbations while holding timestamps correct. A calibration benefit is supported if online estimation approaches the correct-control error and recovers the injected parameter; an ATE gain without parameter recovery would not demonstrate accurate calibration. state-calibrationcloningexperiment-setupoutdoor-atereporting-gaps

Check 2: Separate edge/plane geometry from the number of residuals

Reader-proposed experiment, not a reported reproduction: compare full LIC-Fusion with edge-only, plane-only and no-LiDAR-update variants on identical outdoor and indoor replays. Keep camera tracks, IMU inputs, calibration policy and seeds fixed. For the full, edge-only and plane-only variants, add a residual-budget-matched comparison so a larger measurement count cannot alone explain improvement. Report accepted residual counts and rejection fractions together with outdoor ATE and indoor start-end error. Add full-trajectory indoor ground truth in a new acquisition if available, rather than treating endpoint agreement as path accuracy. If full fusion loses its advantage after matching the budget, the distinct geometry claim is weakened; complementary gains across motion/scene conditions would support it. lidar-geometrylidar-noisevisual-updateoutdoor-ateindoor-protocolindoor-resultsreporting-gaps

8.3 Reading coverage

Visual audit: Visually inspected every supplied PDF page, including the title/authors/affiliations and v2 stamp on p. 1; state and time-offset equations on p. 2; propagation/cloning and LiDAR geometry on p. 3; covariance, gating, visual constraints, QR update, hardware and comparison setup on p. 4; the rig, outdoor trajectories, error plot, Table I and indoor protocol on p. 5; and indoor trajectories, IMU diagnostic, Table II, conclusion and references on p. 6. Figures 1–6 and Tables I–II were visually read. All five final original crops were individually viewed with legible legends, axes and headers. Figure 1 colors match its caption and measurement description; Figure 4 retains the unresolved caption-versus-axis metric conflict. The supplied PDF has no appendix, architecture block diagram or component ablation. Separate supplements and other editions remain outside this pass.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6. Appendix coverage: not present.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • PDF p. 1: title, authors, affiliations, version stamp and abstract
  • PDF pp. 1–2: I. Introduction and Related Work
  • PDF p. 2: II-A. State Vector and II-B. IMU Propagation
  • PDF p. 3: II-C. State Augmentation
  • PDF pp. 3–4: II-D. Measurement Models
  • PDF p. 4: II-E. Measurement Compression and III. Experimental Results setup
  • PDF p. 5: III-A. Outdoor Tests and III-B. Indoor Tests
  • PDF pp. 5–6: IV. Conclusions and Future Work
  • PDF p. 6: References

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • The reviewed artifact is the six-page arXiv:1909.04102v2 dated 1 November 2019. Its exact title and all five authors match the catalog. The catalog separately identifies an IROS 2019 proceedings entry; that edition and earlier arXiv revisions were not supplied, so their differences or equivalence are not established.
  • Text extraction alone omitted figure images; this omission was resolved by visually inspecting all six PDF pages, Figures 1–6, Tables I–II and every final crop.
  • Separate supplemental material availability has not been fully verified; no separate supplement was supplied.
  • Code was not inspected and experiments were not reproduced. No appendix is present in this supplied PDF.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

identityPDF p. 1, title block, author affiliation footnotes and left-margin arXiv stampInspect

Title is LIC-Fusion: LiDAR-Inertial-Camera Odometry; authors are Xingxing Zuo, Patrick Geneva, Woosik Lee, Yong Liu and Guoquan Huang. The stamp reads arXiv:1909.04102v2 [cs.RO], 1 Nov 2019. Affiliations identify Zhejiang University and two University of Delaware departments.

Go to primary source ↓
motivationPDF pp. 1–2, Abstract, I. Introduction and Related Work, contribution bullets; p. 1, Figure 1 and captionInspect

The authors motivate complementary camera, LiDAR and inertial sensing, claim tight single-thread MSCKF fusion, and illustrate red LiDAR edge features, blue plane features, visual tracks and a green estimated trajectory. Optimality is qualified by linearization errors.

Go to primary source ↓
state-calibrationPDF p. 2, II-A, Eqs. (1)–(10) and symbol definitionsInspect

The state contains IMU pose, velocity and biases, camera/LiDAR rigid transforms and clock offsets, and separate pose-clone windows. Equations (7)–(8) define clock correction; JPL quaternion error updates are specified.

Go to primary source ↓
cloningPDF p. 2, II-B, Eqs. (11)–(15); p. 3, II-C, Eqs. (16)–(19)Inspect

IMU kinematics and bias random walks propagate state and covariance. New sensor observations trigger pose cloning at estimated corrected time; the clone Jacobian includes time-offset sensitivity through angular and linear velocity.

Go to primary source ↓
lidar-geometryPDF p. 3, II-D.1, Eqs. (20)–(26) and surrounding feature-association textInspect

High/low curvature selects edge/plane features. Current points are projected to the preceding LiDAR frame and matched with KD-tree indexing. Two old edge points on neighboring rings define a point-to-line residual; pose and extrinsic derivatives enter its linearization.

Go to primary source ↓
lidar-noisePDF p. 4, II-D.1, Eq. (27), Mahalanobis gate and surf-feature paragraphInspect

Distance covariance is propagated from raw point covariances. A chi-squared Mahalanobis test rejects outliers. Three corresponding surf points define a plane; covariance propagation, linearization and gating follow the edge treatment.

Go to primary source ↓
visual-updatePDF p. 4, II-D.2, Eqs. (28)–(30)Inspect

FAST/KLT tracks are triangulated when lost or spanning the window. Reprojection constraints are projected into the feature Jacobian nullspace to remove landmark-position dependence. Camera extrinsic derivatives are nonzero.

Go to primary source ↓
compressionPDF p. 4, II-E, Eqs. (31)–(33)Inspect

LiDAR/visual residuals and Jacobians are stacked with an independent-noise assumption. Givens rotations implement thin QR; the compressed residual and projected noise drive a standard EKF update.

Go to primary source ↓
experiment-setupPDF p. 4, III. Experimental Results, setup and baseline paragraphs; p. 5, Figure 2Inspect

The rig uses an Xsens MTi-300 AHRS, VLP-16 and BFLY-PGE-23S6M global-shutter camera. Extrinsics are initialized offline and refined online. Baselines are the authors’ MSCKF VIO implementation and open-source LOAM; the compared LOAM output matches to a global map, whereas LIC-Fusion has no global map or loop closures.

Go to primary source ↓
outdoor-atePDF p. 5, III-A, Figure 3 and Table I, both rows and all method columnsInspect

An approximately 800 m, four-minute robot route uses RTK GPS and six runs per algorithm to account for RANSAC randomness. Alignment minimizes overall trajectory error. Table I reports average ATEs MSCKF/LIC-Fusion/LOAM = 10.75/4.06/23.08 m and one-sigma variability 3.56/3.42/2.63 m.

Go to primary source ↓
outdoor-ambiguityPDF p. 5, Figure 4 vertical axis and caption; III-A error-reduction paragraph; Table IInspect

The graph axis says Error (m), but caption/text say MSE. The prose reports 2.5 m and 5 m improvements over MSCKF and LOAM, without resolving their relationship to Table I values. The blue trace grows less than the others late in the route; plotted rankings vary over time.

Go to primary source ↓
indoor-protocolPDF p. 5, III-B. Indoor Tests, both columnsInspect

Indoor sequences are handheld at chest height under normal-to-low light and slow-to-aggressive motion. Ground truth is unavailable, so the rig returns to its initial location and start-end error is evaluated. Lighting levels are not assigned individually to the named sequences.

Go to primary source ↓
indoor-resultsPDF p. 6, Table II, Indoor-A/B/C/D rows and all method columns; Figure 5Inspect

Sequence lengths are 39/86/55/189 m. MSCKF errors are 0.99/1.55/49.94/46.03; LIC-Fusion 0.98/1.04/1.55/3.68; LOAM 0.66/0.46/2.44/5.99. Error-column units, indoor repeat counts and uncertainty are not stated. Figure 5 plots estimated trajectories without an indoor ground-truth trace.

Go to primary source ↓
motion-diagnosticPDF p. 5, III-B, Indoor-C description; p. 6, Figure 6, axes, legends and captionInspect

Indoor-C was recorded while vigorously shaking the rig. Figure 6 shows raw angular velocity in rad/s and linear acceleration in m/s² versus time, with separate x/y/z traces. It is a motion-input diagnostic, not an estimation-error or ablation plot.

Go to primary source ↓
future-workPDF pp. 5–6, IV. Conclusions and Future WorkInspect

The authors summarize sparse multimodal odometry and online calibration, and identify efficient integration of LiDAR/camera loop-closure constraints as future work to bound navigation errors.

Go to primary source ↓
reporting-gapsPDF pp. 2–4, II-A–E; pp. 4–6, III-A–B, Tables I–II and Figures 2–6Inspect

The formulation specifies symbolic windows, covariance propagation and gating but no numerical configuration for the listed windows, thresholds or covariances. Experiments give sensor models and aggregate comparisons, without a calibration/sensor/compression ablation, runtime benchmark, processing hardware, software versions, sensor rates or dataset-release instructions.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.