Papers Task Planning
“Task Planning” 태그가 달린 논문 344편 · 필터 해제
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist …
Dataset GenerationDescriptiveMultiple-choiceQuestion Answering+6CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
Large Language Model (LLM) agents can automate cybersecurity tasks and can adapt to the evolving cybersecurity landscape without re-engineering. While LLM agents have demonstrated cybersecurity capabilities on Capture-Th…
Large Language ModelTask PlanningVulnerability DetectionAPEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight
Large Language Models (LLMs) demonstrate strong reasoning and task planning capabilities but remain fundamentally limited in physical interaction modeling. Existing approaches integrate perception via Vision-Language Mod…
Causal InferenceDecision Makingmotion predictionReinforcement Learning (RL)+1Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI Agent
Mobile GUI agents execute user commands by directly interacting with the graphical user interface (GUI) of mobile devices, demonstrating significant potential to enhance user convenience. However, these agents face consi…
Task PlanningREI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. Although recent large language model (LLM)-based task planners achieve amazing …
Large Language ModelRobot Task PlanningTask PlanningLODGE: Joint Hierarchical Task Planning and Learning of Domain Models with Grounded Execution
Large Language Models (LLMs) enable planning from natural language instructions using implicit world knowledge, but often produce flawed plans that require refinement. Instead of directly predicting plans, recent methods…
Robot ManipulationTask PlanningWorld KnowledgeAchieving Scalable Robot Autonomy via neurosymbolic planning using lightweight local LLM
PDDL-based symbolic task planning remains pivotal for robot autonomy yet struggles with dynamic human-robot collaboration due to scalability, re-planning demands, and delayed plan availability. Although a few neurosymbol…
16k8kTask PlanningPIPA: A Unified Evaluation Protocol for Diagnosing Interactive Planning Agents
The growing capabilities of large language models (LLMs) in instruction-following and context-understanding lead to the era of agents with numerous applications. Among these, task planning agents have become especially p…
Instruction FollowingResponse GenerationTask PlanningCoordField: Coordination Field for Agentic UAV Task Allocation In Low-altitude Urban Scenarios
With the increasing demand for heterogeneous Unmanned Aerial Vehicle (UAV) swarms to perform complex tasks in urban environments, system design now faces major challenges, including efficient semantic understanding, flex…
Task PlanningLLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics
We present an embodied robotic system with an LLM-driven agent-orchestration architecture for autonomous household object management. The system integrates memory-augmented task planning, enabling robots to execute high-…
In-Context LearningObjectobject-detectionObject Detection+5Leveraging Pre-trained Large Language Models with Refined Prompting for Online Task and Motion Planning
With the rapid advancement of artificial intelligence, there is an increasing demand for intelligent robots capable of assisting humans in daily tasks and performing complex operations. Such robots not only require task …
Large Language ModelMotion PlanningTask and Motion PlanningTask PlanningNORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
Existing Visual-Language-Action (VLA) models have shown promising performance in zero-shot scenarios, demonstrating impressive task execution and reasoning capabilities. However, a significant challenge arises from the l…
Task PlanningVision-Language-ActionVisual ReasoningEnhancing LLM-Based Agents via Global Planning and Hierarchical Execution
Intelligent agent systems based on Large Language Models (LLMs) have shown great potential in real-world applications. However, existing agent frameworks still face critical limitations in task planning and execution, re…
Task PlanningRobo-Troj: Attacking LLM-based Task Planners
Robots need task planning methods to achieve goals that require more than individual actions. Recently, large language models (LLMs) have demonstrated impressive performance in task planning. LLMs can generate a step-by-…
Backdoor AttackDiversityTask PlanningA Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
Large Language Models (LLMs) exhibit substantial promise in enhancing task-planning capabilities within embodied agents due to their advanced reasoning and comprehension. However, the systemic safety of these agents rema…
BenchmarkingTask PlanningInstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task Planning
Recent advancements in large language models (LLMs) have enabled their use as agents for planning complex tasks. Existing methods typically rely on a thought-action-observation (TAO) process to enhance LLM performance, b…
Meta-LearningMeta Reinforcement LearningRAGreinforcement-learning+4FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment
Geometrically accurate and semantically expressive map representations have proven invaluable to facilitate robust and safe mobile robot navigation and task planning. Nevertheless, real-time, open-vocabulary semantic und…
3D geometryNatural Language QueriesRobot NavigationScene Understanding+1Personality-Driven Decision-Making in LLM-Based Autonomous Agents
The embedding of Large Language Models (LLMs) into autonomous agents is a rapidly developing field which enables dynamic, configurable behaviours without the need for extensive domain-specific training. In our previous w…
Decision MakingSchedulingTask PlanningVisual Environment-Interactive Planning for Embodied Complex-Question Answering
This study focuses on Embodied Complex-Question Answering task, which means the embodied robot need to understand human questions with intricate structures and abstract semantics. The core of this task lies in making app…
Question AnsweringTask PlanningAgent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Computer use agents automate digital tasks by directly interacting with graphical user interfaces (GUIs) on computers and mobile devices, offering significant potential to enhance human productivity by completing an open…
AI AgentTask Planning