Skip to content

Research

Page 62 of 166

GROW$^2$ — Grounding Which and Where for Robot Tool Use

research note

GROW$^2$ — Grounding Which and Where for Robot Tool Use

·8 min read·Yuhong Deng, Yuyao Liu, David Hsu

This paper addresses the challenge of open-world affordance grounding for robot tool use, where a robot must select a suitable tool from a diverse set of objects and precisely localize task-relevan…

researchrobot-tool-useaffordance-groundingvision-language-models3d-reconstruction

Read note → Source paper ↗

MOAR Planner — Multi-Objective and Adaptive Risk-Aware Path Planning for Infrastructure Inspection with a UAV

research note

MOAR Planner — Multi-Objective and Adaptive Risk-Aware Path Planning for Infrastructure Inspection with a UAV

·9 min read·Louis Petit, Alexis Lussier Desbiens

This paper addresses the challenging problem of autonomous UAV navigation for infrastructure inspection, where missions demand safe, energy-efficient, and time-conscious trajectories in dynamically…

researchmulti-objective-path-planningrisk-aware-navigationuav-path-planningadaptive-cost-function

Read note → Source paper ↗

On the Internet, Nobody Knows You're an LLM Bot — Unmasking Web Agents with Multi-Layer Fingerprinting

research note

On the Internet, Nobody Knows You're an LLM Bot — Unmasking Web Agents with Multi-Layer Fingerprinting

·8 min read·Iliana Fayolle, Sihem Bouhenniche, Samuel Pélissier et al.

This paper examines the problem of detecting a new generation of web bots known as LLM-based Web Agents, which leverage large language models (LLMs) combined with browser automation to perform comp…

researchbot-detectionlarge-language-modelsweb-agentsmulti-layer-fingerprinting

Read note → Source paper ↗

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

research note

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

·10 min read·Seongro Yoon, Donghyeon Cho, Jinsun Park et al.

This paper addresses a key limitation in Vision Transformer (ViT) based video models for facial expression recognition (FER) — the tendency of standard self-attention mechanisms to focus on dominant…

researchvideo-transformersfacial-expression-recognitionself-attention-reweightingmasked-video-modeling

Read note → Source paper ↗

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

research note

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

·8 min read·Siddhant Bansal, Zhifan Zhu, Shashank Tripathi et al.

This paper tackles the difficult problem of 3D hand-object pose estimation from challenging in-the-wild egocentric RGB videos, where the hands and objects are heavily occluded and contact regions a…

research3d-hand-pose-estimationegocentric-visionhand-object-interactiontransformer-decoder

Read note → Source paper ↗

Articles are CC BY 4.0 — feel free to quote with attribution