OneCanvas-Qwen3-VL-8B

3D Scene Understanding via Panoramic Reprojection

NeurIPS 2026 (Spotlight)

Paper · Project page · Video · Code

Bartłomiej Baranowski¹, Dave Zhenyu Chen², Matthias Nießner¹

¹Technical University of Munich · ²Huawei

OneCanvas answers questions about a 3D scene. It takes the posed RGB-D views of that scene, places every image patch at the 3D point it was observed from, and spreads those patches over one panoramic canvas seen from a chosen viewpoint. A vision-language model reads that single canvas next to the question and writes the answer in text.

Method

A frozen feature extractor encodes each input view. Depth and camera pose lift every patch feature into 3D, and each lifted patch keeps its own continuous longitude and latitude on an equirectangular canvas built around the chosen origin. Overlapping observations stay separate tokens, so nothing is fused or resampled away. Angular position and source-frame index enter the model through its MRoPE, and a small learned embedding supplies metric 3D coordinates. Only the LoRA adapters and that embedding are trained.

Results

Benchmark Metric OneCanvas-Qwen3-VL-8B
SQA3D EM@1 65.53
VSI-Bench Mean of eight category scores 71.16
SPBench (zero-shot) Mean of SI and MV scores 74.41

These are the scores of the downloadable model on the complete test sets, measured with the commands in Evaluate. Use them to check your setup. Benchmark protocols states the exact aggregation behind each metric, and reproducing any of them requires those settings. The VSI-Bench figure uses gravity-upright ARKitScenes inputs.

Checkpoints

The released model is base Qwen3-VL-8B-Instruct with the stage-1 curriculum adapter and the stage-2 QA adapter merged in, carrying the trained 3D position embedding. The directory is self-contained.

hf download BaranowskiBrt/OneCanvas-Qwen3-VL-8B --local-dir OneCanvas-Qwen3-VL-8B

lineage.json in the downloaded directory records which training checkpoints were merged, in which order, and at which step. resolved_config.json carries the canvas and geometry settings the benchmark runner reads back, and depth_embedding.pt carries the 3D position embedding. All three are required. Loading the weights alone gives a model with a randomly initialized geometry channel and silently wrong metric answers.

The two training stages of the same model are on their own branches.

Branch Contents Use it to
stage1 Stage-1 adapter, 3D position embedding, training config Train your own stage 2
stage2 Stage-2 adapter, 3D position embedding, training config, stage-1 adapter in stage1_lora/ Evaluate the unmerged adapters
hf download BaranowskiBrt/OneCanvas-Qwen3-VL-8B --revision stage1 \
  --local-dir OneCanvas-stage1
hf download BaranowskiBrt/OneCanvas-Qwen3-VL-8B --revision stage2 \
  --local-dir OneCanvas-stage2

The stage-1 adapter applies to the base Qwen3-VL-8B-Instruct, not to the merged model above.

Load the model and answer a question

The model runs through the OneCanvas code, not through a plain AutoModel load. Install it first.

git clone https://github.com/baranowskibrt/onecanvas
cd onecanvas && pip install -e .

This is the shortest path from a downloaded checkpoint to an answer, and the check to run first on a new machine. It needs one scene's posed RGB-D frames, prepared as in Dataset preparation.

from inference import load_onecanvas_model, process_vision_and_generate

CKPT = "OneCanvas-Qwen3-VL-8B"

model, processor = load_onecanvas_model(CKPT)  # includes required geometry state

answer = process_vision_and_generate(
    model, processor,
    question="How many chairs are in this room?",   # given to the model verbatim
    image_paths=frame_paths,   # list of N RGB frames, evenly spaced over the scene
    poses=poses,               # [N, 4, 4] camera-to-world
    depths=depths,             # [N, H, W] MILLIMETRES, sensor-PNG convention
    intrinsics=intrinsics,     # [N, 4] (fx, fy, cx, cy)
    image_dims=image_dims,     # [N, 2] (W, H)
)
print(answer)

The question goes to the model exactly as written, in the format it was trained and benchmarked on. A multiple-choice question lists its options on new lines after the question ("...?\nA. chair\nB. table"), as in VSI-Bench, and a situated question starts with the situation, as in SQA3D ("I am facing the window. What is on my left?").

Feed the frames the way the benchmarks do, or the answers drift:

  • 32 frames spaced evenly over the scene, first and last included, in capture order, and only frames with a valid pose.
  • 640x480 images (320x240 for SQA3D-style questions), with intrinsics in the pixel frame of exactly those images.
  • A Z-up world and OpenCV cameras (x right, y down, z forward), with poses camera-to-world.
  • Gravity-upright images. ARKitScenes stores frames in the sensor's landscape orientation, and the benchmark loader turns them upright first.
  • For a situated question, place the canvas at the agent with center_override=[x, y, z] and yaw_angle= (radians). Otherwise the canvas is centred on the mean camera position, the VSI-Bench setting.

The published numbers use attn_implementation="flash_attention_2". The default sdpa can flip a few answers.

The directory name must contain Qwen3-VL. The benchmark runner routes the model class by substring-matching the path name, and a miss falls through to a different backbone without complaining.

Two failure modes are worth recognizing, because neither raises where it is caused:

  • Skipping load_3d_embeddings leaves the 3D position embedding randomly initialized. The model still answers, and categorical questions look roughly normal, but metric readouts collapse. Measured on the paper checkpoint the cost was 5.4 points of VSI-Bench overall (64.95 against 70.34), concentrated in room size and absolute distance, while SQA3D barely moved. A valid load prints [depth_embed] Loaded depth embedding from ....
  • attn_implementation="flash_attention_2" needs a matching CUDA toolchain. sdpa is the portable choice.

Train

bash training/runs/train_stage1_curriculum.sh

STAGE1_CKPT="$ONECANVAS_CHECKPOINT_DIR/stage1_curriculum/checkpoint-30000" \
  bash training/runs/train_stage2_qa.sh

Stage 1 trains a rank-256 LoRA adapter and the 3D position embedding on the procedural spatial curriculum, with the geometry-to-visual feature norm ratio held at 0.5. Stage 2 merges that adapter into the base model, initializes a fresh rank-64 adapter, and fine-tunes on the scene QA mixture with the geometry scale free. The scripts handle the geometry-scale handoff.

The released model is stage-1 step 30,000 and stage-2 step 7,000 of this procedure. Evaluated as unmerged adapters, the way a new run is scored, it reaches 65.33 SQA3D, 70.81 VSI-Bench and 74.32 SPBench. Reproducing the training procedure gives the configuration, hardware, environment and checkpoint selection. To train only stage 2, start it from the released stage-1 adapter.

STAGE1_CKPT=OneCanvas-stage1 bash training/runs/train_stage2_qa.sh

Evaluate

Reproduce the three released numbers on the downloaded checkpoint. The runner detects a self-contained merged directory, applies its resolved_config.json and skips LoRA loading.

# SQA3D, agent-pose canvas origin, 320x240
python training/run_benchmarks.py --model-path OneCanvas-Qwen3-VL-8B \
  --datasets sqa3d --image-resolution 320x240 --sqa3d-use-agent-pose

# VSI-Bench, 640x480, gravity-upright ARKitScenes inputs
python training/run_benchmarks.py --model-path OneCanvas-Qwen3-VL-8B \
  --datasets vsi_bench --image-resolution 640x480 --upright-arkit

# SPBench (zero-shot), queried-camera canvas origin, 640x480
python training/run_benchmarks.py --model-path OneCanvas-Qwen3-VL-8B \
  --datasets spbench --image-resolution 640x480 --spbench-use-camera-pose
python scripts/spbench_paper_table.py output/<exp>/spbench/qa_results_final.json

training/run_benchmarks_ddp.py takes the same arguments under torchrun and is the practical choice for the full test splits. To evaluate the unmerged adapters from the stage2 branch, pass --from-config OneCanvas-stage2 in place of --model-path. For your own run, pass --stage1-lora and --lora.

Benchmark protocols

Each headline number is a specific aggregation of final_metrics.json, and the other overall-looking fields in that file are different quantities. Using the wrong one moves a result by one to two points without any warning.

Benchmark Headline metric Read it from Inputs
SQA3D EM@1, strict micro-average over all 3,519 test questions final_metrics.json 320x240, canvas origin at the annotated agent pose
VSI-Bench official_overall, the mean of the eight per-category means over 5,130 questions final_metrics.json 640x480, canvas origin at the camera centroid, ARKitScenes frames gravity-upright
SPBench mean of the SI and MV split averages scripts/spbench_paper_table.py 640x480, canvas origin at the queried camera pose
  • On VSI-Bench, all_questions_average is the flat mean over the 5,130 questions and reads about 1.5 points higher than official_overall, because the flat mean overweights the larger categories.
  • On SPBench, official_overall is the flat mean over 1,328 questions, and the benchmark is 1,009 single-image against 319 multi-view, so the flat mean is single-image dominated and reads about 2 points lower than the balanced figure. The two can move in opposite directions.
  • SQA3D per-question-type rows in final_metrics.json are EM@R1, not EM@1. Recompute strict per-type EM@1 from qa_results_final.json when comparing per-type columns against published baselines.
  • All three run on ground-truth sensor depth and poses (--use-gt-all, the default) and 32 evenly spaced frames per scene. ARKitScenes frames are turned gravity-upright by default, whatever the checkpoint's training config says. --no-upright-arkit reads them in the sensor's landscape orientation instead, which costs about a point on VSI-Bench. --scenes-file restricts a run to a listed scene subset.

Evaluation scores are sensitive to the canvas origin, which is why each command names it. Measured on the paper checkpoint, SQA3D scores 65.3 at the annotated agent pose and 61.2 at the scene centroid. SPBench moves by a comparable margin, 69.8 against 62.8 on its flat official_overall, between the queried camera pose and the default anchor.

A merged directory for your own run comes from scripts/export_merged_checkpoint.py, given the stage-1 adapter and the selected stage-2 adapter. Merging rounds weights in bfloat16 and can change generated answers, so evaluate the merged artifact rather than carrying over the scores measured on the adapters.

Citation

If you use OneCanvas in your research, please cite:

@inproceedings{baranowski2026onecanvas,
  title     = {OneCanvas: 3D Scene Understanding via Panoramic Reprojection},
  author    = {Baranowski, Bart{\l}omiej and Chen, Dave Zhenyu and Nie{\ss}ner, Matthias},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}

License

The weights are released under the MIT License, like the code. The base model, Qwen3-VL-8B-Instruct, is released under the Apache-2.0 License.

Downloads last month
53
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BaranowskiBrt/OneCanvas-Qwen3-VL-8B

Finetuned
(616)
this model

Paper for BaranowskiBrt/OneCanvas-Qwen3-VL-8B