Method Detail: VLMOD-GroundNet

Back to Leaderboard
Benchmark: VLMOD
Short name: VLMOD-GroundNet
Long name: VLMOD Grounding Network for Multi-Object Monocular 3D Visual Grounding
Description: VLMOD-GroundNet is an independently implemented and trained neural network for multi-object 3D visual grounding from monocular RGB scenes. The method encodes structured candidate-object information, including object category, appearance, 2D image geometry, 3D dimensions, 3D position, and orientation. Public descriptions are encoded using a learned recurrent text encoder. A learned gated geometry-language fusion module combines object and description representations and predicts three independent grounding logits for every candidate object. The model is trained with scene-level multi-label supervision using binary cross-entropy with logits. The final binary predictions use a threshold selected on a held-out validation split and are exported in the required VLMOD TXT format.
Reference: N/A
Last submitted: September 21, 2026
Published: September 21, 2026 at 01:11:39
Submissions: 1
Project page / code: https://github.com/izharahmaad/VLMOD-GroundNet
Open source: Yes

Benchmark performance

Submission Date F1 (↑) Precision (↑) Recall (↑) TP (↑) FP (↓) FN (↓)
2026-09-21 01:12 63.3937 48.5195 91.4192 2802 2973 263