PAPER REPORTENAll readings ↗

CoPeD-Advancing Multi-Robot Collaborative Perception: A Comprehensive Dataset in Real-World Environments

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Yang Zhou; Long Quang; Carlos Nieto-Granda; Giuseppe Loianno

Affiliations: New York University, Tandon School of Engineering, Brooklyn, NY, USA; U.S. Army Combat Capabilities Development Command, Army Research Laboratory, Adelphi, MD, USA

Source: IEEE Robotics and Automation Letters · ref-1ef0de9829c6511558f2 ↗ · Catalog record

Reading: 459 / 558 · 6 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: CoPeD makes complementary real-world air-ground observations available for collaborative perception, trading inexpensive automatic annotation for unresolved label accuracy and largely qualitative validation. e02-purposee03-platformse04-sensorse09-posese10-annotationse12-depthe13-semantics

At a glanceWhat to know
Research problem
Source description

A collaborator can compensate for an obscured or noisy robot only when its observations contain useful shared scene information. The authors argue that coverage-oriented SLAM datasets and vehicle-centric perception datasets insufficiently capture overlapping views from robots with different mobility, payload and sensing capabilities. CoPeD targets this gap in real environments. e02-purpose

Core mechanism
Source description

A collection spanning three ground robots and two aerial robots overall, with varying team configurations, indoor/outdoor scenes, overlapping viewpoints and asynchronous sensing. e03-platformse05-indoore06-outdoor

A key reported resultCollaborative monocular depth under camera corruption: MultiRobot visually recovers scene structure absent or blurred in Baseline; no numerical improvement is reported.

Qualitative depth-map comparison; no numerical error metric. Figure 7, two shown scenarios with two aerial robots and one ground robot; corrupted RGB in T1 ground column (c) and T2 aerial column (d). Split and sample-selection protocol are not reported.

Referenced graph-network collaboration versus the displayed single-robot baseline. Supports an illustrative robustness use case. The GroundTruth row is source-labeled; its independent measurement provenance is not established, and Section V-B describes model-generated depth annotations. e10-annotationse12-depth

Reading caution
Reader analysis

Foundation-model annotations are predictions, not independently verified human or sensor ground truth. The paper supplies no accuracy audit, class-completeness assessment or temporal-propagation ablation. Comprehensive 2D/3D semantic annotation remains future work. e10-annotationse14-future

Core contributions

  • Source description

    A collection spanning three ground robots and two aerial robots overall, with varying team configurations, indoor/outdoor scenes, overlapping viewpoints and asynchronous sensing. e03-platformse05-indoore06-outdoor

  • Source description

    Raw sensor streams are supplemented by estimated poses and optional zero-shot instance masks, bounding boxes and monocular depth, supporting several perception tasks without requiring one fixed downstream model. e09-posese10-annotations

  • Author claim

    The authors present CoPeD as the first real-world heterogeneous air-ground dataset expressly targeting collaborative perception. The paper demonstrates use cases qualitatively; its historical priority was not independently checked. e02-purposee12-depth

Figure 2. Different payloads and viewpoints are the physical basis of the collection. Original paper, p. 3 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the panels left to right as platform types, not as a network architecture or a complete simultaneous team. The Warthog and Jackal are ground platforms; the ARPL Race robot supplies aerial viewpoints. Section III-A specifies the sensor payloads behind these photographs, including ground LiDAR and aerial forward and downward cameras. The tag board visible on the Jackal connects to another part of the method: downward aerial images support relative localization of the ground robot. The overall collection uses three ground and two aerial robots, while individual sequences use different subsets. Tables II–III, rather than the photographs alone, establish sensor rates and computer specifications. e02-purposee03-platformse04-sensorse06-outdoor

What it supports. The collection deliberately combines mobility and sensing differences. Ground robots support heavier instruments, while aerial robots can observe from above and explore places that constrain ground motion. This physical complementarity motivates cooperative scene understanding, but the photograph itself reports no fusion accuracy or learned-model performance.

Where the evidence stops. Figure 2 documents hardware layout, not a learned architecture. The paper does not provide a new network diagram for its referenced graph-network demonstration; sensor availability and successful cross-robot fusion remain different claims.

2. Motivation

2.1 The problem and the proposed response

Source description

A collaborator can compensate for an obscured or noisy robot only when its observations contain useful shared scene information. The authors argue that coverage-oriented SLAM datasets and vehicle-centric perception datasets insufficiently capture overlapping views from robots with different mobility, payload and sensing capabilities. CoPeD targets this gap in real environments. e02-purpose

2.2 What this reading follows

An aerial camera and a ground robot can observe the same place from very different positions, with different occlusions and sensing capabilities. CoPeD builds a dataset around that complementarity. Its collection spans indoor and outdoor robot teams, combines several sensor rates, and adds estimated poses plus optional foundation-model annotations. Read it as an account of how useful collaborative observations are assembled, then examine how far the demonstrations validate them. The depth examples suggest that information from peers can restore missing structure. They do not supply an error benchmark, and a mismatch between the semantic figure and its caption limits what that example can establish. e02-purposee03-platformse04-sensorse09-posese10-annotationse12-depthe13-semantics

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryDatasets
ArchitectureNot applicable
Prediction paradigmNot applicable
QuadrantNot applicable

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The recorded Datasets category and multisensor collaborative-perception subcategories fit the supplied artifact. Architecture, prediction paradigm and quadrant are appropriately Not applicable: this is a collection/annotation resource, with a referenced perception-network use case, not a proposed joint future/action model. No One Model or inverse-dynamics architecture is established. e02-purposee03-platformse10-annotationse12-depth

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Ground-robot LiDAR, RGB/stereo/depth-camera, IMU, GNSS and wheel-odometry streams; aerial forward/downward cameras, IMU and GPS.
  • Calibration observations and downward-camera images of ground-mounted AprilTag bundles.
  • Raw multirobot sensor sequences; per-robot pose estimates and cross-robot relative transformations.
  • Optional model-generated instance masks, derived 2D boxes and monocular depth annotations; illustrative collaborative perception predictions.

5. Method in detail

5.1 1. Make observations comparable before asking robots to collaborate

Reader analysis

CoPeD starts with a physical design choice: heterogeneous robots should see overlapping parts of the world while contributing different perspectives. The raw observations are only one layer. Their coordinates depend on camera and inertial calibration, the ground multisensor graph, and the board used to initialize alignment across the team. Their timing depends on synchronization across streams that run at different rates. Aerial GPS and OpenVINS, ground inertial and wheel odometry, and downward-camera tag detection supply estimated poses and relative transformations. Reader interpretation: these supporting estimates determine whether two observations can be interpreted as evidence about the same place and time. An apparent perception failure could therefore arise from alignment error as well as image understanding. The paper describes the machinery but does not quantify that uncertainty. e02-purposee03-platformse04-sensorse08-calibratione09-posese11-attributes

Tables II–III. Multirate sensing makes temporal and spatial alignment part of the dataset’s use. Original paper, p. 4 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with the Sensor Rate column, comparing the ground table above with the aerial table below. Ground LiDAR is listed at 10 Hz, the separate ground IMU at 1000 Hz, and ground GNSS at 2 Hz. Aerial GNSS is listed at 8 Hz. The D435i entries separate image channels at 30 Hz from the embedded IMU at 200 Hz. These are different streams, not a single aligned sampling frequency. Then consult the model and resolution columns to identify the observations being synchronized. The computer rows describe onboard acquisition platforms. Section III-A supplies the NTP synchronization arrangement, and Section IV supplies the geometric calibration procedures. e03-platformse04-sensorse08-calibratione11-attributes

What it supports. A timestamped camera frame, inertial sample and LiDAR scan need not represent identical acquisition intervals. The table supports treating synchronization and calibration as prerequisites for interpretation. The listed ground T4 hardware and aerial Jetson Xavier NX characterize the robots; they do not report a training or annotation throughput result.

Where the evidence stops. These are source-reported specifications, not independently measured delivered rates or timing accuracy. Section VI explicitly allows clock synchronization loss from network delay. The table does not show residual timing error or validate every sensor configuration per sequence.

5.2 2. Separate automatic supervision from independent truth

Reader analysis

The optional semantic and depth products are generated by existing models. RAM, Grounding DINO and SAM supply masks, mask post-processing supplies boxes, and ZoeDepth supplies monocular depth. A temporal propagation model extends keyframe segments to neighboring frames. This annotation stage uses zero-shot models rather than a reported CoPeD-specific training procedure. It should also be separated from the graph-network application later in the paper: annotation production and collaborative prediction have different roles. Reader interpretation: a student model evaluated against automatic labels may be learning to agree with the annotation pipeline, which is not necessarily the same as improving physical accuracy. Temporal consistency can similarly reflect persistent error. Figure 6 makes the annotation format clear, but an independent audit would be needed to measure correctness across robot viewpoints and environments. e10-annotationse12-depth

Figure 6. Optional labels combine semantic foundation models, monocular depth and temporal propagation. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Follow the alternating Semantics and Depth rows. The upper pair shows outdoor scenes and the lower pair indoor scenes; the column groups distinguish aerial and ground viewpoints. Colored overlays and boxes illustrate the semantic annotation products, while grayscale panels illustrate depth. Section V-B identifies RAM, Grounding DINO and SAM as the mask-generation components, with boxes obtained by post-processing masks. ZoeDepth supplies monocular depth annotations. The same section describes propagation of keyframe segments to consecutive frames. That temporal procedure is specified in the text, not demonstrated as an on/off comparison in this still-image grid. The paper treats these high-level annotations as optional additions to raw data. e10-annotations

What it supports. CoPeD supplies more than camera recordings: the examples demonstrate the form of automatically produced scene labels for several robot views. Their role is to make downstream perception experiments easier to construct. The figure establishes annotation availability in example frames, while leaving the accuracy and completeness of those annotations unmeasured.

Where the evidence stops. Model-generated masks and depth are not independently certified targets. No metric depth scale or accuracy score appears here, and individual overlay scores do not establish calibration. Temporal propagation could preserve an incorrect label as well as a correct one.

5.3 3. Read recovery as a hypothesis about useful peer information

Reader analysis

Figure 7 provides a concrete way to reason about collaboration. When one robot’s RGB input is corrupted, its standalone depth estimate loses scene structure; the collaborative output recovers some structure while other robots retain usable views. Section VII attributes the information exchange to feature-map communication through a referenced graph neural network. Reader interpretation: correctly matched peer observations are a plausible cause of recovery, but the figure alone cannot isolate that cause from training or model-capacity differences. It supplies neither a quantitative test split nor a matched ablation protocol. Figure 8 adds a separate caution: its semantic caption and grayscale output panels disagree. The defensible reading is therefore a qualitative depth use case and an unresolved semantic demonstration, not a measured general robustness guarantee or a result about executed robot actions. e12-depthe13-semantics

5.4 Training and inference

During training

Source description

Annotation is zero-shot: the source says the models are not trained on task-specific data. It specifies no CoPeD training objective, optimization schedule, checkpoint versions or annotation-validation protocol. The separate graph-network use case likewise lacks a training recipe; task-specific training details cannot be inferred from the cited method alone. e10-annotationse12-depthe15-access

During inference

Source description

The demonstrated collaborative application exchanges feature maps among two aerial robots and one ground robot using a referenced graph neural network. It predicts depth and, according to the authors, semantics under sensor noise. No future-state rollout, action decoder or closed-loop control evaluation is defined. e12-depthe13-semantics

5.5 Implementation flow

  1. Collect complementary views

    Warthog and Jackal platforms carry heavier sensing payloads, while ARPL aerial robots supply elevated and downward views. Teams split into aerial-ground pairs; HOUSEB changes pairings and includes an outdoor-to-indoor transition. Single-aerial-robot sequences provide additional samples. These collection behaviors create overlap and occlusion variation rather than a learned action policy. e03-platformse06-outdoore11-attributes

  2. Align asynchronous measurements

    ROS 1 robots synchronize to a laptop clock using NTP. Reported rates differ substantially: ground LiDAR 10 Hz, ground IMU 1000 Hz, ground GNSS 2 Hz, aerial GNSS 8 Hz and D435i images 30 Hz. Kalibr handles aerial calibration; a multisensor graph rooted at the ground OS1 LiDAR and a shared calibration board establish spatial relationships. e03-platformse04-sensorse08-calibration

  3. Estimate poses

    Aerial GPS is fused with OpenVINS stereo visual-inertial odometry; ground estimation uses IMU and wheel odometry. Downward cameras detect calibrated eight-tag bundles, with PnP supporting inter-robot localization. These outputs are estimated poses: no independent trajectory-error evaluation establishes ground-truth accuracy. e03-platformse09-poses

  4. Generate optional annotations

    RAM, Grounding DINO and SAM supply instance masks, subsequently converted to bounding boxes. ZoeDepth produces monocular depth. Keyframe segments propagate to consecutive frames using the method cited as Track Anything. The authors favor monocular depth over classical short-baseline stereo, but provide no numerical comparison validating that choice. e10-annotations

6. Experiments & results

CoPeD records real indoor and outdoor air-ground robot teams with overlapping camera and LiDAR views, pose estimates, and optional automatic perception labels. Its contribution is a heterogeneous collection and annotation pipeline. A graph-network demonstration suggests qualitative depth recovery under camera corruption; it does not establish numerical benchmark gains or action-execution performance.

Source and visual limitations
Reader analysis

The paper supplies platform photographs, sensor and sequence tables, annotation examples and qualitative perception demonstrations. It contains no proposed neural-architecture diagram, quantitative perception-results table or controlled ablation. Figure 2 therefore illustrates the physical setup, Tables II–IV document specifications and collection statistics, and Figure 7 serves as a qualitative corruption diagnostic rather than a measured ablation. Figure 8’s semantic caption conflicts with its depth-like outputs. The source does not give displayed method equations, per-task accuracy metrics or measured action-execution results. e03-platformse04-sensorse07-sequencese10-annotationse12-depthe13-semantics

6.1 Read the original evidence

Table IV. The source quantifies collection entries, without defining a benchmark split. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each row as one named entry in the source table. Keep Time, RGB Frames and Instances/Frame separate: none is a perception-accuracy metric. NYUARPL is listed as 400 seconds with 48,000 RGB frames and 9 instances/frame; HOUSEA and HOUSEB each list 440 seconds and 52,800 frames, but their last-column values differ. The horizontal divider separates the three indoor rows from the three outdoor rows. Section III gives the environment and team descriptions; Section VI adds variations in formation and single-robot sequences. The table does not explain how frames are aggregated across robots or how instances/frame is calculated. e05-indoore06-outdoore07-sequencese11-attributes

What it supports. The rows establish that the illustrated collection contains several environments and different scene compositions. HALLWAY lists 3 instances/frame, whereas CHAIR-1 lists 11. Those source statistics can guide questions about scene diversity, but they do not by themselves establish class balance, annotation accuracy or generalization.

Where the evidence stops. The prose describes four indoor and five outdoor sequences, beyond these six rows, and mentions additional single-robot data. No mapping reconciles that inventory. Do not treat the table as an exhaustive split or infer an overall dataset total from it.

Figure 8. The published semantic example has an unresolved mismatch between image content and caption. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Begin with the printed row labels: RGB, Baseline and MultiRobot. The top row shows color images with box or mask overlays. The lower rows are smooth grayscale maps that resemble the T2 depth outputs in Figure 7. Now compare the caption, which calls this semantic estimation and describes the rows as ground-truth semantics, a single-robot baseline and collaboration. Section VII also claims recovery of houses in semantic segmentation. Those descriptions do not establish why the lower rows look like depth maps. Preserve the original visual and read the claim as an author statement whose intended semantic comparison cannot be recovered unambiguously from this artifact. e12-depthe13-semantics

What it supports. The upper row shows annotated house regions across three viewpoints, while the lower rows show a visible difference between baseline and collaborative outputs. The source does not provide the legend or corrected panel needed to identify that difference as improved semantic segmentation. A numerical semantic conclusion is therefore unsupported.

Where the evidence stops. The caption says panel (a)’s sensor is corrupted, yet its displayed RGB overlay is not the noise image used in Figure 7. The intended input/output correspondence remains unresolved. This edition retains the source discrepancy rather than assigning new meanings to the rows.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Recorded sequence coverage

Table IV lists six named entries; these are collection statistics, not a train/test benchmark.

NYUARPL: 400 s; 48,000; 9. CHAIR-1: 180 s; 10,800; 11. HALLWAY: 200 s; 12,000; 3. HOUSEA: 440 s; 52,800; 8. HOUSEB: 440 s; 52,800; 9. FOREST: 225 s; 27,000; 5.

Duration; RGB frames; instances/frame, as reported

Three indoor and three outdoor table rows; no performance baseline applies.

The table documents scene variation. Frame aggregation and instances/frame are not defined, and the rows do not reconcile the larger sequence counts in the prose. e05-indoore06-outdoore07-sequences

Collaborative monocular depth under camera corruption

Figure 7, two shown scenarios with two aerial robots and one ground robot; corrupted RGB in T1 ground column (c) and T2 aerial column (d). Split and sample-selection protocol are not reported.

MultiRobot visually recovers scene structure absent or blurred in Baseline; no numerical improvement is reported.

Qualitative depth-map comparison; no numerical error metric

Referenced graph-network collaboration versus the displayed single-robot baseline.

Supports an illustrative robustness use case. The GroundTruth row is source-labeled; its independent measurement provenance is not established, and Section V-B describes model-generated depth annotations. e10-annotationse12-depth

Collaborative semantic perception

Section VII and Figure 8 claim a three-robot semantic demonstration with corruption in aerial panel (a).

The authors report recovered houses; the published visual does not unambiguously substantiate a semantic comparison.

Author-described qualitative robustness; no segmentation score

Rows labeled Baseline and MultiRobot; both appear as grayscale depth-like maps beneath an RGB overlay row.

Retain the author claim with an unresolved graphic/caption discrepancy. No quantitative or visually unambiguous semantic gain can be extracted. e13-semantics

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure 7. A qualitative corruption diagnostic suggests useful information transfer between robots. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. First partition the columns into T1, panels (a)–(c), and T2, panels (d)–(f). Each scenario contains two aerial views and one ground view. The caption identifies corrupted sensors in (c) and (d), matching the noise-filled RGB panels. Read vertically within each column: source-labeled GroundTruth, single-robot Baseline, then MultiRobot. The corrupted columns are especially informative because their baseline maps lose structure while the collaborative maps recover visible scene boundaries or objects. Section VII attributes this demonstration to a referenced graph neural network that communicates feature maps. The graphic supplies qualitative input/output comparisons; it contains no communication arrows, architectural internals or numeric error scale. e10-annotationse12-depth

What it supports. The depicted collaborative maps restore tree or house-area structure that is absent or blurred in the single-robot maps, including the corrupted views. This is an illustrative sign that peer observations can be useful. It is not an estimate of average improvement over a specified dataset or corruption distribution.

Where the evidence stops. This is a qualitative diagnostic, not an isolated ablation with documented training controls. The source gives no numerical errors or corruption recipe. Its GroundTruth label does not establish independent measurement provenance; Section V-B describes automatic depth annotations.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Foundation-model annotations are predictions, not independently verified human or sensor ground truth. The paper supplies no accuracy audit, class-completeness assessment or temporal-propagation ablation. Comprehensive 2D/3D semantic annotation remains future work. e10-annotationse14-future

Reader analysis

Collection inventory is incomplete: prose describes four indoor and five outdoor sequences, versus six Table IV rows, and further single-robot sequences lack counts. Section III-C describes both one-pair cases and all outdoor runs as two subteams. Its ground speed up to 1.5 m/s and Section V-A maximum 1.75 m/s are not reconciled. e05-indoore06-outdoore07-sequencese09-posese11-attributes

Reader analysis

Selected qualitative demonstrations do not quantify robustness, generalization, communication cost or real-time performance. Figure 8 also conflicts with its semantic caption. Neither action execution nor benefits of changing formation are experimentally measured here. e12-depthe13-semanticse06-outdoor

Reader analysis

Realistic disturbances include GPS denial/drift, changing IMU behavior and network-related clock loss. Their presence is useful for stress testing, but no error distributions or calibration/pose uncertainty are reported. e08-calibratione09-posese11-attributes

7.2 Questions for discussion

  1. Would collaboration still improve depth on independently measured targets when the corrupted robot and corruption severity change?
  2. How much of annotation consistency comes from correct temporal propagation, and how much could be persistent labeling error?
  3. Which sequence identities and split boundaries reconcile the paper’s table with its collection narrative?

8. Reproducibility audit

8.1 Requirements and known gaps

Source description

Reconstruction requires sequence inventories, robot identities, timestamps, camera/IMU/LiDAR calibration and tag-bundle geometry. The paper names ROS 1, Kalibr, OpenVINS and annotation components, but leaves exact software versions, checkpoints, keyframe selection and synchronization residuals unspecified. Listed onboard computers describe acquisition equipment, not a training-compute budget. e03-platformse04-sensorse08-calibratione09-posese10-annotations

Reader analysis

The paper links a dataset and video but states no dataset license or official split. A proposed evaluation should separate scenes and contiguous trajectories, prevent synchronized cross-robot views from leaking across splits, and audit labels independently before treating them as evaluation targets. e07-sequencese10-annotationse11-attributese15-access

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Does recovery depend on correctly matched peer features?

Reader-proposed check, not an experiment performed here: on a held-out complete trajectory with independently validated depth targets, compare the same collaborative model using correct peer features, zeroed peer features, and time-shuffled peer features. Keep target frames, input corruption, model weights and evaluation masks fixed; also report a separately trained single-robot baseline with its training budget disclosed. Corrupt each robot in turn and report error on valid target pixels separately for corrupted and intact views. The hypothesis predicts greater recovery with correctly matched peers than with absent or mismatched peers. Equal gains under shuffled peers would weaken the information-transfer explanation. Record synchronization and pose residuals because correspondence errors could otherwise confound the comparison. e03-platformse08-calibratione09-posese11-attributese12-depth

Check 2: Does propagation improve correctness or only temporal smoothness?

Reader-proposed check, not an experiment performed here: select separated indoor and outdoor clips spanning both robot viewpoints, including an occlusion and the HOUSEB transition. Independently annotate instance masks on sampled frames. Compare framewise zero-shot masks with keyframe-propagated masks while holding source models, prompts and sampled frames fixed; document keyframe spacing as an experimental choice because the paper omits it. Measure mask overlap with the independent annotations, missed instances and temporal identity consistency, reporting failures by scene and viewpoint. Propagation should improve consistency without systematically reducing correctness. Smoother but less accurate masks would falsify the assumption that consistency alone is useful denoising. Keep every synchronized robot view from a trajectory in the same evaluation partition. e07-sequencese10-annotationse11-attributese14-future

8.3 Reading coverage

Visual audit: All eight pages were rendered and visually inspected, including the title/authors/version and affiliations on p. 1; dataset taxonomy on p. 2; platform layout and synchronization description on p. 3; sensor specifications and indoor examples on p. 4; sequence statistics, outdoor examples, calibration and AprilTag pose illustration on p. 5; annotation panels and procedures on p. 6; depth and semantic demonstrations on p. 7; and conclusion and component references on p. 8. Figures 1–8 and Tables I–IV were inspected. All six final crops were individually viewed. Figure 7 corruption positions agree with its caption; Figure 8’s row/content/caption mismatch is disclosed. No separate appendix is present. Linked dataset files, code, video and separate supplements remain outside this reading.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8. Appendix coverage: not present.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • PDF p. 1: title, abstract, supplementary-material pointers and author metadata
  • Sections I–II: Introduction and Related Works, pp. 1–3
  • Section III-A: Robots and Equipment, pp. 3–4
  • Sections III-B/C: Indoor and Outdoor Sequences, pp. 4–5
  • Sections IV and V-A: Calibration and Pose Estimation, p. 5
  • Section V-B: Zero-shot Semantics and Depth Annotation, p. 6
  • Section VI: Dataset Attributes, pp. 6–7
  • Sections VII–VIII: Use-Cases, Applications and Conclusion, pp. 7–8
  • References [1]–[37], p. 8

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • The extraction-only image gap was resolved by visually inspecting all eight PDF pages, Figures 1–8 and Tables I–IV, and all six final crops. All four supplied text chunks were read individually.
  • Identity and edition scope: the observed title and all four authors match the catalog. The supplied artifact is arXiv:2405.14731v1, 23 May 2024, labeled an accepted RA-L preprint. The catalog cites the journal edition, volume 9(7), pages 6416–6423; this report reviews the supplied preprint pages 1–8. The journal artifact and any revision differences were not supplied or compared.
  • No separate supplement or appendix was supplied; the PDF has no appendix. The linked video, dataset files, repository and code were not inspected; no experiments were reproduced.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e01-identityPDF p. 1, title, author/affiliation block, arXiv margin stamp and acceptance footnoteInspect

The title and four authors match the catalog: Yang Zhou, Long Quang, Carlos Nieto-Granda, and Giuseppe Loianno. The artifact is arXiv:2405.14731v1, dated 23 May 2024, an IEEE Robotics and Automation Letters preprint accepted 21 May 2024. Affiliations are New York University, Tandon School of Engineering, and the U.S. Army Combat Capabilities Development Command, Army Research Laboratory.

Go to primary source ↓
e02-purposePDF pp. 1–3, Abstract, Sections I–II, Figure 1 and Table I, CoPeD rowInspect

The authors motivate real-world heterogeneous air-ground collaborative perception with complementary, overlapping sensor views. They distinguish this goal from coverage-oriented SLAM datasets and vehicle-centric perception; priority and novelty statements are author claims.

Go to primary source ↓
e03-platformsPDF p. 3, Section III-A and Figure 2; p. 4, Section III-A continuationInspect

Collection uses three ground robots and two aerial robots overall: Warthog and Jackal ground platforms and ARPL aerial robots. Ground platforms carry LiDAR, inertial, GNSS and camera sensors; aerial robots carry forward D435i, downward IMX219, PX4 and GPS. ROS 1 and NTP synchronization to a laptop over a shared Wi-Fi subnetwork are specified. Downward-camera AprilTag detection and PnP provide relative localization, using an eight-tag bundle.

Go to primary source ↓
e04-sensorsPDF p. 4, Tables II–III, LiDAR, IMU, GNSS, RGBD and computer rowsInspect

Reported rates include ground LiDAR 10 Hz, ground Microstrain IMU 1000 Hz, ground GNSS 2 Hz, aerial GNSS 8 Hz, D435i RGB/stereo 30 Hz and embedded IMU 200 Hz. Ground computers list i7-8700 CPUs, 32 GB RAM and NVIDIA T4 GPUs with 16 GB RAM; the aerial computer is a Jetson Xavier NX with 8 GB RAM. These are platform specifications, not measured annotation or training costs.

Go to primary source ↓
e05-indoorPDF p. 4, Section III-B and Figure 3Inspect

The text describes three indoor sequences with one Warthog and one aerial robot, plus one with two ground and two aerial robots. NYUARPL occupies a 38 m by 60 m indoor space, with lab and hallway subteams; aerial robots fly 2.0 m above ground robots, which move at 0.5 m/s.

Go to primary source ↓
e06-outdoorPDF p. 5, Section III-C and Figure 4Inspect

The text describes three outdoor sequences with two Warthogs and two aerial robots, plus two with one Warthog and one aerial robot. HOUSE spans 150 m by 30 m; FOREST spans 100 m by 80 m. Aerial robots fly at 2.0–10.0 m and ground robots drive up to 1.5 m/s in this collection description. HOUSEB switches aerial-ground pairings. The same paragraph describes all outdoor sequences as two subteams, leaving the one-pair cases unreconciled.

Go to primary source ↓
e07-sequencesPDF p. 5, Table IV, all six rows and five column headersInspect

The table lists NYUARPL: Indoor, 400 s, 48000 RGB frames, 9 instances/frame; CHAIR-1: Indoor, 180 s, 10800, 11; HALLWAY: Indoor, 200 s, 12000, 3; HOUSEA: Outdoor, 440 s, 52800, 8; HOUSEB: Outdoor, 440 s, 52800, 9; FOREST: Outdoor, 225 s, 27000, 5. It does not define frame aggregation or instances/frame, nor reconcile its six entries with the larger sequence counts in Sections III-B/C.

Go to primary source ↓
e08-calibrationPDF p. 5, Section IV, Aerial Robot Calibration, Ground Robot Calibration and Air-Ground CalibrationInspect

Aerial calibration uses Kalibr for camera intrinsics and camera–IMU extrinsics, with IMU Allan covariance analysis. Ground calibration uses a multisensor graph with Ouster OS1 as reference. A board viewed by all forward-facing cameras initializes team extrinsics. Numerical calibration residuals are not reported.

Go to primary source ↓
e09-posesPDF p. 5, Section V-A and Figure 5Inspect

Aerial pose estimation combines GPS with OpenVINS stereo VIO; ground estimation uses IMU and wheel odometry. Calibrated AprilTag bundles observed from downward aerial cameras provide relative transformations between bundle and camera frames. The section lists maximum aerial/ground speeds of 2.1/1.75 m/s, without reconciling the ground maximum with Section III-C. No trajectory-error benchmark or independently measured pose ground truth is supplied.

Go to primary source ↓
e10-annotationsPDF p. 6, Section V-B and Figure 6; p. 8, References [31]–[37]Inspect

RAM, Grounding DINO and SAM produce instance masks; boxes are obtained by mask post-processing. ZoeDepth generates monocular depth labels. The authors prefer it to classical stereo for their short-baseline cameras and state that their grayscale cameras preclude the learning-based stereo methods they consider. Track Anything, cited as [37], propagates keyframe segments to consecutive frames. These are zero-shot annotations without task-specific training; annotation accuracy and propagation ablations are not quantified.

Go to primary source ↓
e11-attributesPDF pp. 6–7, Section VIInspect

HOUSEB includes an outdoor-to-indoor transition by race5. Team formation, overlap and occlusion vary. Listed disturbances include camera color jitter, temporally discontinuous LiDAR intensities, denied or drifting GPS, temperature-related IMU changes and network-related clock synchronization loss. Additional single-aerial-robot HOUSE and FOREST sequences are described without an enumerated inventory.

Go to primary source ↓
e12-depthPDF p. 7, Section VII and Figure 7, rows RGB/GroundTruth/Baseline/MultiRobot and columns (a)–(f)Inspect

The demonstration uses feature-map communication through the referenced graph neural network with two aerial robots and one ground robot. T1 occupies columns (a)–(c), T2 columns (d)–(f); RGB sensors (c) and (d) are corrupted. MultiRobot depth maps visually recover structure relative to the single-robot Baseline. The source supplies no numerical metric, split, sample count, uncertainty or corruption-generation protocol for this demonstration.

Go to primary source ↓
e13-semanticsPDF p. 7, Figure 8 graphic, Figure 8 caption and Section VII; compare Figure 7 columns (d)–(f)Inspect

Section VII claims improved semantic robustness and recovery of houses. Figure 8 is captioned as semantics estimation with a corrupted image sensor in (a), but its top row is labeled RGB and shows colored box/mask overlays, while Baseline and MultiRobot rows are grayscale maps resembling the T2 depth panels in Figure 7. No semantic color legend or quantitative segmentation result resolves this graphic/caption mismatch.

Go to primary source ↓
e14-futurePDF pp. 7–8, Section VIII, especially p. 8 future-work paragraphInspect

The authors plan more comprehensive 2D/3D semantic annotations, more environments and robot configurations, and possible real-to-sim use. These are future directions, not demonstrated evaluation results.

Go to primary source ↓
e15-accessPDF p. 1, Supplementary Material; pp. 3–7, Sections III–VIIInspect

The paper points to a dataset repository and supplementary video. The supplied paper does not specify a dataset license, a versioned file inventory or official train/validation/test splits. Platform, calibration and annotation methods are described, but the collaborative demonstration lacks an optimization recipe, model checkpoints, training compute and evaluation metrics.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.