| Description: |
We propose a cross-modal matching network for multi-object 3D visual grounding from monocular RGB images. Given an image, three textual descriptions and a set of detected 3D objects, the method predicts for each object which descriptions it matches (a multi-label binary classification over the three descriptions).
The method consists of three components:
(1) Text encoding. The three textual descriptions are encoded by a pre-trained language model (MiniLM) into semantic sentence embeddings via mean pooling, which are then projected into a common embedding space.
(2) Geometry encoding with image-level 3D features. Each object's 16-field structured features (2D bounding box, 3D size, 3D location, orientation, truncation, occlusion, type and color) are normalized and encoded by a multi-layer perceptron with learnable type/color embeddings. In parallel, a convolutional backbone (ResNet-18) extracts image-level 3D features by applying RoIAlign to each object's 2D region, which are fused with the structured geometry features to enhance the object representation.
(3) Cross-attention fusion. Object tokens serve as queries and text tokens serve as keys/values in a cross-attention module, followed by a matching head that produces multi-label (N x 3) binary logits indicating which of the three descriptions each object matches.
The model is trained with a binary cross-entropy loss over the multi-label matching matrix. |