computer vision
A field of AI focused on enabling machines to interpret and understand visual information from the world, including tasks such as image recognition and processing.
- Bringing SAM to new heights: leveraging elevation data for tree crown segmentation from drone imagery
- DAMamba: Vision State Space Model with Dynamic Adaptive Scan
- Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image Restoration
- Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task Conditioning
- Equivariance by Contrast: Identifiable Equivariant Embeddings from Unlabeled Finite Group Actions
- FedRTS: Federated Robust Pruning via Combinatorial Thompson Sampling
- From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- Latency NMS Attacks: Is It Real Life or Is It Just Fantasy?
- Learning Shared Representations from Unpaired Data
- LuxDiT: Lighting Estimation with Video Diffusion Transformer
- MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
- More effort is needed to protect pedestrian privacy in the era of AI
- Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
- Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- Robust Hyperbolic Learning with Curvature-Aware Optimization
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- SegMASt3R: Geometry Grounded Segment Matching
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
- Toward Human Deictic Gesture Target Estimation
- Towards Generalizable 3D Human Pose Estimation via Ensembles on Flat Loss Landscapes
- VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
- Web-Scale Collection of Video Data for 4D Animal Reconstruction