| Description: |
VLMOD-GroundNet is an independently implemented and trained neural network for multi-object 3D visual grounding from monocular RGB scenes. The method encodes structured candidate-object information, including object category, appearance, 2D image geometry, 3D dimensions, 3D position, and orientation. Public descriptions are encoded using a learned recurrent text encoder. A learned gated geometry-language fusion module combines object and description representations and predicts three independent grounding logits for every candidate object. The model is trained with scene-level multi-label supervision using binary cross-entropy with logits. The final binary predictions use a threshold selected on a held-out validation split and are exported in the required VLMOD TXT format. |