zero-shot generalization
Zero-shot generalization is the ability of a model to make predictions about classes or tasks that it has never encountered during training. This capability is particularly important in environments where training data may be scarce or where flexibility is required.
- $\texttt{G1}$: Teaching LLMs to Reason on Graphs with Reinforcement Learning
- Anchored Diffusion Language Model
- Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement Learning
- CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
- Domain Adaptive Hashing Retrieval via VLM Assisted Pseudo-Labeling and Dual Space Adaptation
- DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization
- HyPINO: Multi-Physics Neural Operators via HyperPINNs and the Method of Manufactured Solutions
- IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
- Know Thyself by Knowing Others: Learning Neuron Identity from Population Context
- Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
- MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- Multi-dataset Joint Pre-training of Emotional EEG Enables Generalizable Affective Computing
- NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG Signals
- Native-Resolution Image Synthesis
- OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild
- One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
- Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
- Repurposing Marigold for Zero-Shot Metric Depth Estimation via Defocus Blur Cues
- SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models
- Self supervised learning for in vivo localization of microelectrode arrays using raw local field potential
- SensorLM: Learning the Language of Wearable Sensors
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
- Tree-Guided Diffusion Planner
- UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss
- VeriThinker: Learning to Verify Makes Reasoning Model Efficient
- Zero-Shot Trajectory Planning for Signal Temporal Logic Tasks