← 전체 목록

2026-07-28 AI 리서치 브리핑

최신 VLM, sLLM, on-device AI 논문과 연구 블로그를 한눈에 정리합니다. 중복 기사 방지를 위해 URL 기준으로 추적합니다.

총 10건 요약 자동 생성

VLM 업데이트

멀티모달 비전-언어 모델의 최신 논문과 리더보드 변화

Twins: Learn to Predict Unified Representations with Focal Loss

Paper arXiv cs.CV (recent)

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations -- semantic features (e.g., ViT) for understanding and low-level latents (e.g., VAE) for synthesis -- resulting in mismatched latent spaces. We propose Twins, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase. However, jointly modeling Twins in a Diffusion Transformer exposes a severe optimization imbalance: the model fits the ViT component well but struggles to match the VAE latent distribution. We trace this imbalance to three sources of heterogeneity: frequency bias, intrinsic dimensionality, and condition-aligned vs condition-independent uncertainty. To address it, we adapt a focal regression objective for flow matching that upweights large-error VAE dimensions, better balancing optimization across the ViT and VAE components. On ImageNet, this yields up to 10.57 gFID gain over naive MSE loss without classifier-free guidance. Twins also performs competitively on multimodal understanding benchmarks and improves reconstruction fidelity, narrowing the gap between understanding- and generation-oriented representations.

원문 보기

CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation

Paper arXiv cs.CV (recent)

자율 주행에서 충돌 예측은 정확한 조기 경고뿐만 아니라 위험 요인에 대한 해석 가능한 추론을 요구하지만, 기존 방법들은 불투명하거나 동적 장면에 부적합했습니다. 본 논문은 개념 인식 위험 주의(CARA)라는 해석 가능한 시공간 프레임워크를 제안하며, 이는 사고 서술에서 도메인 기반 위험 개념을 도출하고 이를 비디오 프레임과 정렬하여 개념 궤적을 구성합니다. CARA는 이러한 개념 궤적을 통해 모델이 어디에 주의하고 위험을 어떻게 예측하는지에 직접적인 영향을 미쳐 예측 정확도와 경고 조기성을 향상시키면서 의미론적으로 근거 있는 개념 증거를 제공합니다.

English Collision anticipation in autonomous driving requires accurate early warnings and interpretable reasoning about risk factors, which existing methods often lack due to opacity or unsuitability for dynamic scenes. We propose CARA (Concept-Aware Risk Attention), an intrinsically interpretable spatio-temporal framework that derives domain-grounded risk concepts from accident narratives. CARA aligns these concepts with video frames to form evolving concept trajectories, which then guide spatial attention, temporal attention, and anticipation. This approach improves anticipation accuracy and warning earliness while providing sparse, semantically grounded concept evidence.

원문 보기

AI 뉴스 & 리서치

기업/연구기관의 주요 발표와 블로그 업데이트

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Paper Hugging Face Papers

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Paper Hugging Face Papers

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Paper Hugging Face Papers

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Paper Hugging Face Papers

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Interactive Training 2: Auditable Control Plane for Live Model Training

Paper Hugging Face Papers

Interactive Training 2: Auditable Control Plane for Live Model Training에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Robot-Factored World Models via Robot Rendering

Paper arXiv cs.CV (recent)

Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two distinct processes: they are first realized into robot motion by the robot body and controller, and the scene then responds through contact and object motion. Conditioning directly on action commands asks the world model to learn the realization process itself, while conditioning on logged future states leaks the interaction outcomes it is meant to predict. We propose robot-factored world models, which move two robot-specific factors outside the world model. First, action realization: each command is rolled through the robot's own controller and kinematics into a deployment-available nominal trajectory, a middle signal that avoids both action-realization learning and future-state leakage. Second, robot rendering: this nominal trajectory is rendered through the robot URDF, factoring the robot's geometry, kinematics, and appearance out of the model and into explicit rendered robot geometry. To resolve depth ambiguity, we pair end-effector depth with scene depth, giving geometric cues for contact and occlusion beyond image-plane overlap. Together, camera-aware static RGB/depth context and rendered robot geometry form a shared visual world-model interface that stays consistent across viewpoints and robot embodiments, so the model sees the action only as visible robot geometry and learns how objects respond to it. Our experiments show that the rendered interface outperforms vector-conditioned baselines and generalizes to unseen robot embodiments at inference. We further demonstrate that our model generates robot manipulation videos from human demonstrations by retargeting and rendering the hand motion as robot geometry.

원문 보기

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Paper arXiv cs.CV (recent)

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstruction and structured motion perception. SM4RT introduces Structure-of-Motion to represent scene dynamics, where scene motion is decomposed into a compact set of motion bases, each represented as a temporal sequence of 6D twists in SE(3). Dense scene motion is then recovered by sparse, time-shared per-pixel assignment weights over these bases, ensuring points on the same object share a common rigid-body motion trajectory. SM4RT introduces a parallel motion geometry encoder and decoder that jointly infer 3D geometry, world-coordinate motion, and scene kinematic structure in a single forward pass from monocular RGB video. SM4RT achieves strong motion reconstruction performance while preserving the geometric structure of scene motion.

원문 보기

CARE: Anti-entanglement Ultrasound Image Segmentation via Channel-Aware Region Extrication

Paper arXiv cs.CV (recent)

초음파 이미지 분할은 병변과 주변 조직, 아티팩트가 시각적으로 얽혀 정확도가 저해되는 문제가 있습니다. 기존 방법들은 특징 추출이나 컨텍스트 통합에 집중하여 모호한 예측에 취약했지만, CARE는 채널 인식 영역 추출을 통해 병변 증거를 시각적으로 얽힌 컨텍스트에서 점진적으로 분리합니다. 이 방법은 학습된 표현에서 목표-컨텍스트 구별을 촉진하여 우수한 성능을 달성하며, 초음파 분할의 시각적 모호성을 해결하는 효과적인 솔루션임을 입증합니다.

English Accurate ultrasound image segmentation is challenged by target-context entanglement, where lesion cues mix with surrounding tissues and artifacts. Existing methods often struggle with ambiguous predictions, focusing on feature extraction or context aggregation rather than explicitly distinguishing cues. We propose CARE (Channel-Aware Region Extrication), a framework that progressively extricates lesion evidence from visually entangled context by separating encoded responses based on lesion relevance. This approach promotes target-context discrimination in learned representations, achieving superior performance and validating representation extrication as an effective solution for ultrasound segmentation's inherent visual ambiguity.

원문 보기

참고한 소스