feilong tang
- Decoding Causal Structure: End-to-End Mediation Pathways Inference
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
- UniViT: Unifying Image and Video Understanding in One Vision Encoder