Generalized Predictive Model for Autonomous Driving
1. Paper overview
In one sentence: GenAD learns driving-video dynamics around a frozen image backbone, enabling controllable visual futures and a small planning head while retaining costly pretraining and unproven closed-loop performance. e-data-scalee-temporale-video-objectivee-extensionse-plane-training
| At a glance | What to know |
|---|---|
| Research problem | Source description Driving models trained in narrow geographic and sensor settings can fail under new environments and intentions. The authors propose video prediction as a scalable way to learn driving representations, but fast camera motion and independently moving agents make temporal correspondence difficult. Their problem combines obtaining diverse data, preserving future consistency and transferring learned features to downstream tasks. e-probleme-temporal |
| Core mechanism | Source description OpenDV-2K combines 1747 hours of YouTube driving with 312 hours from seven public datasets: 2059 hours and 65.1M front-view frames. Its geographic counts are title-derived estimates, not verified position annotations. e-data-scale |
| A key reported result | nuScenes future-video prediction: 15.4; 184 FID ↓; FVD ↓. Table 2; GenAD trained on OpenDV-2K, including nuScenes. Exact evaluation sample counts and split identifiers are not provided in this PDF. GenAD-nus: 15.4/244; DriveGAN: 73.4/502; DriveDreamer: 52.6/452; DrivingDiffusion: 15.8/332. Best reported FVD, but training data and conditioning differ. DriveDreamer and DrivingDiffusion use 3D layouts; DrivingDiffusion is not evaluated as future prediction. This is not zero-shot nuScenes evidence. e-data-scalee-video-results |
| Reading caution | Author claim The authors acknowledge that model capacity impedes training efficiency and real-time deployment. e-limits |
Core contributions
- Source description
OpenDV-2K combines 1747 hours of YouTube driving with 312 hours from seven public datasets: 2059 hours and 65.1M front-view frames. Its geographic counts are title-derived estimates, not verified position annotations. e-data-scale
- Source description
GenAD separates driving-image adaptation from temporal learning, inserting causal temporal and decoupled spatial attention around a frozen image backbone. Extensions test trajectory-controlled video generation and a learned waypoint head. e-imagee-temporale-extensions
Table 1. A much larger driving corpus, with inclusion and geographic-estimation caveats visible. Original paper, p. 3 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Start with the leftmost marks: a check means the source contributes to OpenDV-2K, while a cross identifies a comparison dataset outside that mixture. Then read the final two rows together. OpenDV-YouTube contributes 1747 hours and 60.2M frames; the complete mixture reaches 2059 hours and 65.1M frames. The dagger on country and city counts matters: these are GPT estimates from video titles. The sensor column also changes from fixed setups to uncalibrated cameras. That broadens appearance variation without providing a common calibrated geometric interface for every example. e-data-scalee-data-curatione-language-data
What it supports. The dataset contribution is both scale and heterogeneity. YouTube supplies most of the video hours, while the incorporated public datasets add other sensor and language sources. nuScenes is explicitly included, so the later nuScenes generation table should not be presented as a zero-shot evaluation of GenAD.
Where the evidence stops. Hours and geographic coverage do not directly measure balanced traffic behavior or annotation accuracy. Several table entries are unspecified, and the geographic estimates are not GPS-verified counts. Uploader-level validation separation is described separately.
2. Motivation
2.1 The problem and the proposed response
Driving models trained in narrow geographic and sensor settings can fail under new environments and intentions. The authors propose video prediction as a scalable way to learn driving representations, but fast camera motion and independently moving agents make temporal correspondence difficult. Their problem combines obtaining diverse data, preserving future consistency and transferring learned features to downstream tasks. e-probleme-temporal
2.2 What this reading follows
The paper builds a driving predictor in two complementary ways: it expands the available experience with OpenDV-2K, then changes how an image diffusion model exchanges information across frames. This edition follows that information flow from web-video annotations to causal temporal blocks, before separating two downstream uses. In one, a supplied trajectory controls an imagined video; in the other, frozen features feed a waypoint predictor. The tables make the boundaries especially clear: better video scores do not establish safe driving, and a cheap planning adaptation does not erase the cost of learning the underlying representation. e-data-scalee-temporale-video-objectivee-extensionse-plane-training
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Foundational work |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
This table preserves the labels recorded at reading time. The current major category is WAMs. View the current classification.
3.1 Evidence-based assessment
Supports the recorded classification
The recorded foundational-work classification, with dataset and neural-world-simulator subcategories, fits the combined data and predictive-model contribution. Leaving the action-model architecture and prediction quadrant not applicable is defensible: future video is predicted jointly across frames, not jointly with action outputs. Trajectories are conditions in GenAD-act, while a separately trained MLP produces waypoints from frozen features. The IDM serves evaluation, so its presence does not make the control pathway inverse dynamics. e-data-scalee-video-objectivee-extensionse-action
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
4.2 Equations and their role
5. Method in detail
5.1 Learn appearance before asking the model to predict motion
Start with what is supervised. OpenDV-2K pairs driving frames with a scene context and an ego command. YouTube contexts come from BLIP-2, whereas commands come from a video classifier and a wording dictionary; public datasets contribute additional descriptions and logged trajectories. Stage one uses the concatenated text condition to train the SDXL UNet to remove noise from individual driving-image latents. The text encoders and autoencoder remain frozen. Stage two preserves that adapted image model and adds temporal blocks. Only future latents are corrupted, and only their noise predictions enter the video loss; the clean historical latents supply observations. This staged construction explains why good appearance and temporal consistency are evaluated separately: the method deliberately learns them through different parameter updates. e-language-datae-imagee-video-objective
Figure 3. Driving appearance is adapted first; temporal reasoning is then learned around frozen layers. Original paper, p. 5 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read panel (a) left to right: image-domain transfer precedes video prediction. In panel (b), blue temporal reasoning blocks are inserted before the original spatial attention, conditional cross-attention and feed-forward layers. Snowflakes mark inherited frozen components during stage two. Panel (c) flows upward: causal temporal attention is followed by two decoupled spatial-attention layers. In its bottom illustration, the orange query at the middle time attends to itself and the blue earlier location; the dark-gray later location is masked. The zero-initialized final layers sit inside residual branches. These markers agree with the caption and the stage-two objective. e-imagee-temporale-video-objective
What it supports. The design separates appearance knowledge from temporal adaptation. Clean past latents condition noisy future latents, and training updates the added temporal blocks rather than the inherited image parameters. The spatial layers expand how information can reach temporal attention when camera motion shifts content across image locations.
Where the evidence stops. The mask constrains feature access in the depicted model. It is not evidence that the learned representation recovers physical causality. Joint denoising here concerns future frames; the diagram does not show joint prediction of future images and control actions.
5.2 Follow information through the causal and spatial layers
A temporal attention layer at a fixed image grid location has a difficult job when the ego vehicle moves: the same object may occupy a different location in the next frame. GenAD therefore follows causal temporal attention with attention along the two spatial axes, and interleaves these blocks repeatedly among the inherited transformer layers. Relative temporal position embeddings distinguish frames, while zero initialization of each added block's final layer limits disruption when training begins. The causal mask restricts a query to itself and earlier times, as Figure 3 illustrates. Table 3 then tests the design cumulatively. Its mixed FID and FVD changes matter: the source's argument for improved consistency should be read alongside the complete metric pattern and the selected artifact examples in Figure 6. e-temporale-ablation
5.3 Trace the different roles played by trajectories
Reader analysis: the easiest way to avoid overinterpreting GenAD is to track whether a trajectory enters or leaves each pathway. For action-conditioned prediction, the trajectory enters through a Fourier embedding and cross-attention; the output is an imagined video. The inverse dynamics model then reads that video for evaluation, yielding a consistency error rather than executing a maneuver. For planning, historical frames enter a frozen encoder and an MLP outputs waypoints. The described planner does not select actions by searching generated futures. Table 5 therefore supports representation transfer to an open-loop prediction task, while Table 4 supports a degree of action controllability in simulation. These are complementary demonstrations, but their combination does not by itself establish a unified, closed-loop world-action controller. e-extensionse-actione-plane-video-objective
5.4 Training and inference
During training
Stage one runs 300K iterations on 32 NVIDIA Tesla A100 GPUs with total batch 256. Stage two runs 112.5K iterations on 64 GPUs with total batch 64; the second-stage GPU model is not explicitly restated. Both use 256 × 448 frames; video clips span four seconds at 2 Hz. Text is dropped with probability 0.1. e-training
Stage two freezes the inherited parameters and optimizes only new temporal blocks. Historical latents remain clean, future latents receive diffusion noise, and only future outputs contribute to the noise-prediction loss. e-video-objective
During inference
Given clean observations and text, iterative diffusion denoising produces future latents and hence images; classifier-free guidance is enabled by training-time text dropout. The supplied main text does not specify the sampler, number of denoising steps or guidance scale. e-imagee-video-objectivee-traininge-detail-gaps
The action-conditioned demonstration uses two starting frames and six trajectory waypoints to generate six future frames. Its inverse dynamics model is an evaluator that maps generated videos back to trajectories, not the planner or a demonstrated vehicle controller. e-actione-extensions
5.5 Implementation flow
- Build multimodal training examples
The YouTube collection contains 2139 videos from 43 uploaders, with three uploaders reserved for validation. BLIP-2 describes frames; a classifier trained on Honda-HDD-Action assigns 14 command categories to four-second clips, then a dictionary diversifies wording. Public-dataset text is expanded with GPT and logged trajectories supply commands. e-data-curatione-language-data
- Adapt spatial appearance
Initialize the latent UNet from SDXL and fine-tune it on driving image-text pairs. CLIP text encoders and the autoencoder stay frozen. This stage learns the driving image distribution before temporal modules are introduced. e-image
- Connect past and future features
Each new block applies causal attention along time, then separate horizontal and vertical attention. The mask permits the query's own frame and earlier frames. Blocks appear before inherited spatial attention, cross-attention and feed-forward layers; zero-initialized final layers stabilize insertion. e-temporal
- Separate simulation from planning
GenAD-act supplies Fourier-embedded, linearly projected future trajectories through cross-attention. Planning instead feeds frozen encoder features from two historical frames to an MLP. The supplied planning method does not describe sampling candidate futures, scoring them or executing feedback control. e-extensionse-plan
6. Experiments & results
GenAD adapts an image diffusion model into a driving-video predictor using OpenDV-2K and trainable temporal reasoning blocks. Its future-video generator can receive trajectory conditions, while a separate lightweight planner reads frozen encoder features. The evidence supports visual prediction and economical planning adaptation, with substantial pretraining and limited deployment validation.
6.1 Read the original evidence
Table 2. The strongest FVD comes from broad driving pretraining, under unequal comparison conditions. Original paper, p. 6 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. First inspect the training-dataset column and the prediction flag, before comparing scores. The main GenAD uses OpenDV-2K, whereas GenAD-nus and the listed baselines use nuScenes. Both GenAD variants score 15.4 in FID, but their FVD values differ: 184 for the full mixture and 244 for nuScenes-only training. The asterisks identify methods requiring 3D layouts. DrivingDiffusion also carries a cross in the prediction column, so it is not evaluated as future prediction in this table. Those qualifications remain visible in the original notes and must travel with the numerical comparison. e-video-resultse-data-scalee-training
What it supports. GenAD obtains the lowest reported FVD and matches its nuScenes-only variant in FID. This supports a benefit in the paper's video-quality evaluation from the full training setup. The table alone does not identify whether additional data volume, data diversity or another training difference causes that benefit.
Where the evidence stops. Inputs and training data are not matched across all rows. The supplied main PDF omits exact evaluation sample counts and uncertainty estimates. FID and FVD characterize generated imagery and video, not driving safety or closed-loop control success.
Figure 7. Changing a supplied trajectory changes the imagined ego motion. Original paper, p. 8 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow each row from the starting view toward the right along the time arrow. The four rows share the intersection but supply different intended trajectories, drawn as yellow dots in the first column. Compare the building and tree positions in the final column: the different viewpoint changes illustrate stop, left-turn, forward and right-turn conditions. The text explains how these conditions enter the model: Fourier trajectory features are projected and supplied through cross-attention. The method describes two starting frames and six future waypoints producing six future frames; the figure displays selected frames rather than every model input and output. e-extensionse-action
What it supports. The examples support qualitative controllability by a supplied ego trajectory. The separate Table 4 on this inspected page supplies a quantitative proxy: action prediction error falls from 2.54 for text-only GenAD to 2.02 for GenAD-act. An inverse dynamics evaluator obtains those errors by reading the generated video back into a trajectory.
Where the evidence stops. The yellow dots are conditioning inputs, not a trajectory chosen by the generator. The evaluator also gives ground-truth video a nonzero error of 0.90. These images and proxy scores therefore do not demonstrate physically executed maneuvers.
Table 5. A small waypoint head trades some planning accuracy for inexpensive adaptation. Original paper, p. 8 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read trainable parameter count alongside both error columns. GenAD reports 0.8M trainable parameters with ADE 1.23 and FDE 2.31. It improves on ST-P3's 2.65 and 3.73 but remains behind UniAD's 1.03 and 1.65. The asterisks mark multi-view baselines, while GenAD uses the front view alone. The caption says evaluation protocols align with UniAD, yet that does not make input information identical. Section 3.3 explains the actual pathway: two historical frames pass through a frozen GenAD encoder, and a learned MLP predicts future waypoints. This table evaluates that adaptation. e-plane-extensionse-training
What it supports. The predictive encoder contains features useful enough for a small supervised head to produce competitive open-loop trajectories. With pre-extracted features, the authors report ten minutes of planning adaptation on one V100. That result concerns fitting the downstream planner; it is separate from training the driving-video model.
Where the evidence stops. The column counts trainable parameters, not the total deployed network. Foundation pretraining and feature-extraction costs must be considered separately. Open-loop displacement errors do not establish collision rates or feedback stability, and the table does not label metric units.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| nuScenes future-video prediction Table 2; GenAD trained on OpenDV-2K, including nuScenes. Exact evaluation sample counts and split identifiers are not provided in this PDF. | 15.4; 184 FID ↓; FVD ↓ | GenAD-nus: 15.4/244; DriveGAN: 73.4/502; DriveDreamer: 52.6/452; DrivingDiffusion: 15.8/332. Best reported FVD, but training data and conditioning differ. DriveDreamer and DrivingDiffusion use 3D layouts; DrivingDiffusion is not evaluated as future prediction. This is not zero-shot nuScenes evidence. e-data-scalee-video-results |
| Temporal design ablation YouTube evaluation; cumulative variants trained on an OpenDV-2K subset for 75K steps. | Baseline 18.32/244.44/0.8405; +deep interaction 17.96/201.69/0.8409; +causality 16.54/207.45/0.8550; +spatial attention 17.67/189.54/0.8652. FID ↓ / FVD ↓ / CLIPSIM ↑ | Final variant versus plain temporal-attention baseline. Final FVD and CLIPSIM improve, but causality worsens FVD relative to the preceding row and spatial attention worsens FID. Sequential additions do not isolate every interaction. e-ablation |
| Action-conditioned video consistency GenAD-act fine-tuned on nuScenes; IDM translates predicted video into a trajectory for comparison with the supplied future trajectory. | 2.02 Action Prediction Error: trajectory L2 distance ↓; units and precise aggregation unspecified. | Text-only GenAD 2.54; ground-truth video 0.90. Trajectory conditioning improves the reported proxy. The nonzero ground-truth error demonstrates evaluator imperfection; the result measures simulation consistency, not executed driving. e-action |
| Open-loop nuScenes planning Frozen GenAD encoder and trainable MLP; front-view inputs; evaluation protocols aligned with UniAD. | 1.23 / 2.31; 0.8M trainable parameters. ADE ↓ / FDE ↓; units not labeled in Table 5. | Multi-view UniAD: 1.03/1.65, 58.8M; multi-view ST-P3: 2.65/3.73, 10.9M. The small head beats ST-P3 but trails UniAD. Trainable parameter counts exclude frozen-backbone cost. The reported ten-minute adaptation on one V100 uses pre-extracted features and excludes foundation-model pretraining. e-plane-training |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Table 3. Temporal design improves the final video metrics, but individual changes expose tradeoffs. Original paper, p. 8 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the rows as cumulative additions, not independent removals from the final model. The baseline uses plain temporal attention. Deep interaction inserts the new blocks more frequently among inherited layers, followed by temporal causality and then decoupled spatial attention. Lower FID and FVD are preferred; higher CLIPSIM is preferred. Notice where the bold entries split: the causality row has the best FID, while the final row has the best FVD and CLIPSIM. Section 4.2 specifies that these variants were trained for 75K steps on a subset, so these scores belong to a different experiment from Table 2. e-ablatione-temporal
What it supports. Deep interaction lowers FVD from 244.44 to 201.69. Adding causality raises CLIPSIM from 0.8409 to 0.8550 but also raises FVD to 207.45. The final spatial addition reaches FVD 189.54 and CLIPSIM 0.8652 while worsening FID from 16.54 to 17.67. The improvements are metric-dependent.
Where the evidence stops. This sequential study does not isolate all component interactions, and no seed variability is reported. The authors argue that the regressions do not faithfully represent visual-quality decline; that interpretation is distinct from the numerical changes themselves.
7. Analysis & limitations
7.1 What the evidence leaves open
The authors acknowledge that model capacity impedes training efficiency and real-time deployment. e-limits
Zero-shot transfer and language following are illustrated qualitatively. Table 2 measures generation quality, and Table 5 measures open-loop waypoint error; none establishes collision avoidance, closed-loop stability or physical deployment performance. e-zero-shote-languagee-video-resultse-plan
Uploader separation and described geofencing support held-out evaluation, but exact geographic filtering is absent. BLIP-2 contexts describe static frames, and geographic coverage is estimated from video titles. Neither scale nor these annotations establish uniformly reliable dynamics supervision. e-data-curatione-language-datae-data-scalee-zero-shot
7.2 Questions for discussion
- How much zero-shot improvement comes from additional hours, geographic diversity or language annotations?
- Would video-pretrained encoder features outperform image-stage features under identical planning supervision and head capacity?
8. Reproducibility audit
8.1 Requirements and known gaps
Reproduction needs the curated OpenDV-2K mixture, uploader/geographic partitions, command classifier and dictionary, text-generation procedures, SDXL initialization and added attention blocks. The supplied text gives trimming and filtering procedures but not exact trimming lengths, complete keyword lists, optimizer settings, learning rates or full sampling configuration; appendix pointers leave these unresolved here. e-data-curatione-language-datae-imagee-temporale-traininge-detail-gaps
A proposed mechanism check should toggle causality and decoupled spatial attention under matched data and training budgets, reporting FID, FVD and CLIPSIM together. A second check should compare the same waypoint head on frozen image-stage versus video-stage encoders to test whether temporal pretraining adds planning value. e-ablatione-imagee-video-objectivee-plan
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Separate the causal mask from the spatial-attention benefit
Reader-proposed, not run: keep deep interaction fixed and train a two-by-two comparison with causal versus bidirectional temporal attention and decoupled spatial attention present versus absent. Use the same image-stage checkpoint, OpenDV-2K subset, clip sampling, 75K-step budget and evaluation clips, with several matched seeds. Report FID, FVD and the paper's CLIPSIM once its omitted implementation details are recovered; also compare frame consistency under large viewpoint shifts. A consistent causality benefit across both spatial settings would support an independent mask effect. A benefit only when spatial attention is present would indicate an interaction that the cumulative table cannot isolate. e-temporale-ablatione-traininge-detail-gaps
Check 2: Test whether video pretraining improves the planning representation
Reader-proposed, not run: train the same 0.8M-parameter waypoint MLP on frozen stage-one image features and frozen stage-two video features. Use identical two-frame front-view observations, nuScenes splits, target waypoints, head initialization, optimization and UniAD-aligned evaluation protocol. Match the feature interface and report any required adapter. Compare ADE/FDE over multiple seeds, and separately measure feature extraction, head fitting and inference costs. If the video features reliably lower both errors, temporal pretraining adds planning value beyond image adaptation. If performance is indistinguishable, Table 5 alone cannot attribute the useful representation specifically to video prediction. e-imagee-video-objectivee-extensionse-plan
8.3 Reading coverage
Visual audit: Visually inspected the title, all seven figures, all five tables and claim-supporting method, training, evaluation and limitation text on PDF pages 1–8. Checked Figure 3's flow direction, causal mask, frozen-layer symbols and zero-initialized branches against its caption and Equations (1)–(2); no claim-relevant conflict was found. Viewed every final crop. Table 1's inclusion/estimate notes, Table 2's prediction/layout notes and Table 5's input/protocol notes remain in the crops. Pages 9–11 contain references and were read in the text chunks only. The main PDF references appendices that were not supplied; no supplementary images, code or experimental runs were inspected.
PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8. Appendix coverage: not present.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Title, affiliations, version watermark and abstract (PDF p. 1)
- 1. Introduction (PDF pp. 1–3)
- 2. OpenDV-2K Dataset, including 2.1–2.2 (PDF pp. 3–4)
- 3. GenAD Framework, including 3.1–3.3 and Equations (1)–(2) (PDF pp. 4–6)
- 4. Experiments, including 4.1–4.3 (PDF pp. 6–8)
- 5. Limitations and Discussion (PDF p. 8)
- References (PDF pp. 9–11)
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Separate supplemental material availability has not been fully verified.
- Identity/version: the observed title and all 14 authors match the catalog. This is the CVPR 2024 CVF open-access accepted version, proceedings pages 14662–14672. Its watermark states that only the watermark differs from the accepted version. The final IEEE proceedings artifact and other revisions were not supplied or compared; no arXiv revision is established.
- Text extraction does not reconstruct figure images; this limitation was addressed by inspecting PDF pages 1–8 and all six final crops.
- Separate supplemental material availability has not been fully verified. The supplied PDF has no appendix; cited Appendices C.1.1, C.1.2 and D and additional appendix visualizations were not supplied or read.
- All five text chunks, including references, were read individually. Reference pages 9–11 were read as text, without a separate visual pass.
- Code, linked resources and dataset files were not inspected, and no experiments were reproduced.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e-identityPDF p. 1 (proceedings p. 14662), title, author block and CVF watermark
The title and all 14 authors match the catalog. Five affiliation entries are printed. The watermark identifies the CVPR open-access version as identical to the accepted version except for the watermark and points to IEEE Xplore for the final proceedings version.
Go to primary source ↓e-problemPDF p. 2 (14663), Section 1, motivation and three research questions
The paper targets generalization across environments and driving intentions through scalable video data, predictive modeling of dynamic scenes, and downstream adaptation.
Go to primary source ↓e-data-scalePDF p. 3 (14664), Table 1 including footnotes, Figure 2 and Section 2
OpenDV-2K contains 2059 hours and 65.1M front-view frames; 1747 hours are from YouTube and 312 hours from public datasets. Table 1 marks included sources and reports at least 40 countries and 244 cities, estimated by GPT from video titles; YouTube cameras are uncalibrated.
Go to primary source ↓e-data-curationPDF p. 3 (14664), Section 2.2, Driving Video Collection and Curation
The authors select 2139 videos from 43 uploaders, reserve all videos from three uploaders for validation, trim beginnings and endings, and remove unsuitable frames through keyword-assisted manual checking of BLIP-2 descriptions.
Go to primary source ↓e-language-dataPDF p. 4 (14665), Section 2.2, language annotation and public-dataset integration
Commands come from a Honda-HDD-Action video classifier with 14 action categories over four-second sequences and a free-form expression dictionary. BLIP-2 supplies frame contexts. Public-dataset contexts are expanded with GPT and commands inferred from logged trajectories.
Go to primary source ↓e-imagePDF p. 4 (14665), Section 3.1 and Equation (1)
SDXL supplies a latent denoising UNet. Stage one fine-tunes all UNet parameters on driving image-text pairs while freezing the CLIP text encoders and autoencoder. The objective predicts Gaussian noise conditioned on concatenated context and command.
Go to primary source ↓e-temporalPDF p. 5 (14666), Figure 3(a–c), caption and Section 3.2
Temporal reasoning blocks are interleaved before frozen spatial attention, conditional cross-attention and feed-forward layers. Each block uses causal temporal attention followed by horizontal and vertical attention; the orange query attends to itself and blue locations while the future dark-gray location is masked. New final layers are zero-initialized; temporal relative position embeddings distinguish frames.
Go to primary source ↓e-video-objectivePDF pp. 5–6 (14666–14667), Section 3.2 Training and Equation (2)
A clip is partitioned into clean historical latents and noisy future latents. The loss uses only future-frame noise predictions. The stage-one parameters theta are frozen and the inserted temporal-block parameters phi are trained.
Go to primary source ↓e-trainingPDF p. 6 (14667), Section 4.1 Setup and Protocols
Stage one uses 300K iterations, 32 NVIDIA Tesla A100 GPUs and total batch 256. Stage two uses four-second clips at 2 Hz, 112.5K iterations, 64 GPUs and total batch 64. Both resize frames to 256 by 448 and drop text with probability 0.1 for classifier-free guidance. The second-stage GPU model is not explicitly restated.
Go to primary source ↓e-extensionsPDF p. 6 (14667), Section 3.3, Action-conditioned Prediction and Planning
Action conditioning embeds a future ego trajectory using Fourier features and a linear layer, then supplies it through conditional cross-attention. Planning instead passes frozen UNet-encoder features from two historical frames to a learned MLP predicting waypoints.
Go to primary source ↓e-video-resultsPDF p. 6 (14667), Table 2, its footnotes and Section 4.2
GenAD reports nuScenes FID 15.4 and FVD 184; GenAD-nus reports 15.4 and 244. DriveGAN reports 73.4/502, DriveDreamer 52.6/452, and DrivingDiffusion 15.8/332. The latter two use 3D layouts; DrivingDiffusion is not evaluated as future prediction in this table.
Go to primary source ↓e-zero-shotPDF pp. 6–7 (14667–14668), Section 4.2 and Figure 4 with caption
Qualitative zero-shot comparisons use held-out geofenced OpenDV-YouTube scenarios, Waymo, KITTI and Cityscapes. Blue boxes mark generated futures. The authors compare GenAD to I2VGen-XL, VideoCrafter1 and DMVFN, and discuss weaker transfer from the nuScenes-only variant.
Go to primary source ↓e-languagePDF p. 7 (14668), Figure 5 and Section 4.2
Three language conditions applied to the same rainy-intersection starting frames produce different illustrated futures. These are qualitative examples, without a reported language-following success rate.
Go to primary source ↓e-ablationPDF p. 7 (14668), Section 4.2 Ablation Study; PDF p. 8 (14669), Table 3 and Figure 6
Variants train on an OpenDV-2K subset for 75K steps. Baseline FID/FVD/CLIPSIM is 18.32/244.44/0.8405; adding deep interaction gives 17.96/201.69/0.8409; adding temporal causality gives 16.54/207.45/0.8550; adding decoupled spatial attention gives 17.67/189.54/0.8652. Figure 6 illustrates corresponding artifact and consistency changes.
Go to primary source ↓e-actionPDF pp. 7–8 (14668–14669), Section 4.3 Action-conditioned Prediction; PDF p. 8, Figure 7 and Table 4
GenAD-act uses two starting frames and six future waypoints to predict six future frames. A nuScenes inverse dynamics model recovers a trajectory from video for L2 comparison to the conditioning trajectory. Table 4 reports errors 2.54 for GenAD, 2.02 for GenAD-act and 0.90 for ground-truth video; the text reports a 20.4% reduction.
Go to primary source ↓e-planPDF p. 6 (14667), Section 3.3 Planning; PDF p. 8 (14669), Table 5 with caption and Section 4.3 Planning Results
Open-loop nuScenes planning reports GenAD ADE/FDE 1.23/2.31 with 0.8M trainable parameters, ST-P3 2.65/3.73 with 10.9M, and UniAD 1.03/1.65 with 58.8M. GenAD uses front-view images; starred baselines use multiple views. Protocols align with UniAD. With pre-extracted encoder features, planner adaptation is reported to take ten minutes on one NVIDIA Tesla V100.
Go to primary source ↓e-limitsPDF p. 8 (14669), Section 5 Limitations and Discussion
The authors identify increased model capacity as a challenge for training efficiency and real-time deployment and propose distillation and broader downstream reuse as future directions.
Go to primary source ↓e-detail-gapsPDF p. 4 (14665), Section 2.2 appendix pointers; PDF p. 6 (14667), Section 4.1 appendix pointer; PDF pp. 7–8, Figures 4 and 7 captions
Dataset construction and annotation details are assigned to Appendices C.1.1 and C.1.2; additional training and sampling details to Appendix D; further visualizations to an appendix. These appendices are not part of the supplied eleven-page main-paper PDF.
Go to primary source ↓8.5 Primary sources
Generalized Predictive Model for Autonomous Driving ↗
PDF · 7,123 extracted words
Source fingerprint
90ae20a2ea7d04e5eee4485316b4389b59c73ee2dc5cd93c27bde150cea243ad