Multimodal 3D representation learning

The representation learning the other two threads are built on.

The problem. Scientific data is structured, sparse, expensive to label and captured through several complementary sensors or assays at once. Learning useful representations from it means confronting all four properties together, whether the object is a molecule or a street.

What I build. Methods for learning from structured 3D data under weak supervision. (Yin et al., 2022) pre-trains detectors on unlabelled point clouds by contrasting region proposals; (Yin et al., 2022) and (Wang et al., 2023) push the same objective into the semi-supervised and domain-adaptive settings, where labels exist but not for the distribution you care about; (Li et al., 2023) supervises segmentation from cheaper modalities instead of dense masks. On the multimodal side, (Yin et al., 2024) fuses camera and LiDAR at both the instance and the scene level, and (Yin et al., 2023) models temporal structure with graph message passing and spatiotemporal attention.

Where it is going. These techniques were developed for perception, and the transfer to molecular systems is direct: pre-training when labels are scarce, adapting across distributions, and fusing modalities that each see part of the object. That transfer is what made the move into protein design a continuation rather than a restart.

References

2024

  1. IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
    Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li, and 3 more authors
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

2023

  1. SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point Cloud
    Yan Wang, Junbo Yin, Wei Li, Pascal Frossard, and 2 more authors
    In AAAI Conference on Artificial Intelligence (AAAI), 2023
    * Equal contribution (Y. Wang, J. Yin).
  2. LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous Driving
    Xiang Li, Junbo Yin, Botian Shi, Yikang Li, and 2 more authors
    In AAAI Conference on Artificial Intelligence (AAAI), 2023
    * Equal contribution (X. Li, J. Yin).
  3. Graph Neural Network and Spatiotemporal Transformer Attention for 3D Video Object Detection from Point Clouds
    Junbo Yin, Jianbing Shen, Xin Gao, David Crandall, and 1 more author
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

2022

  1. ProposalContrast: Unsupervised Pre-training for LiDAR-based 3D Object Detection
    Junbo Yin, Dingfu Zhou, Liangjun Zhang, Jin Fang, and 3 more authors
    In European Conference on Computer Vision (ECCV), 2022
  2. Semi-supervised 3D Object Detection with Proficient Teachers
    Junbo Yin, Jin Fang, Dingfu Zhou, Liangjun Zhang, and 3 more authors
    In European Conference on Computer Vision (ECCV), 2022