PTG

The Perceptually-enabled Task Guidance (PTG) program, funded by the Defense Advanced Research Projects Agency (DARPA), aims to develop artificial intelligence (AI) technologies to help users perform complex physical tasks while making them more versatile by expanding their skillset and more proficient by reducing their errors. PTG seeks to develop methods, techniques, and technology for artificially intelligent assistants that provide just-in-time visual and audio feedback to help with task execution. The goal is to provide users of PTG assistants with wearable sensors (head-mounted cameras and microphones) that allow the assistant to see what they see and hear and what they hear, and augmented reality (AR) headsets that allow assistants to provide feedback through speech and aligned graphics. The target assistants will learn about tasks relevant to the user by ingesting knowledge from checklists, illustrated manuals, training videos, and other sources of information. They will then combine this task knowledge with a perceptual model of the environment to support mixed-initiative and task-focused user dialogs. The dialogs will assist a user in completing a task, identifying and correcting an error during a task, and instructing them through a new task, taking into consideration the user’s level of expertise.
New York University (NYU) is one of the teams participating in the project, driving advancements by developing innovative solutions.

Data Provenance and Analytics


Illustration of the research project
ARGUS: Augmented Reality Guidance and User-modeling System

ARGUS enables the interactive exploration and debugging of all components of the data ecosystem needed to support intelligent task guidance. ARGUS has two operation modes: “Online” (during task performance), and “Offline” (after performance). Users can use these two modes separately if needed, for instance, to perform real-time debugging through the online mode. In another usage scenario, users may start by using the online mode to record a session and then explore and analyze the data in detail using the offline mode.

Illustration of the research project
HuBar: A Visual Analytics Tool to Explore Human Behaviour based on fNIRS in AR guidance systems

To effectively model performer behavior, we must determine a method of summarizing and comparing performer behavior across sessions. This necessitates a meaningful way to compare multimodal time series data (e.g., gaze origin and direction, acceleration, angular velocity, fNIRS sensor readings) of different durations. HuBar provides a visual analytics tool for summarizing and comparing task performance sessions in AR, highlighting correlations between cognitive workload and performer motion data.

Illustration of the research project
ARPOV: Expanding Visualization of Object Detection in AR with Panoramic Mosaic Stitching

Inspired by the Panorama View of ARGUS, we built ARPOV: a standalone visual analytics tool that enables troubleshooting of object detection results, with features tailored to analyzing object detection (ground truth and predicted bounding boxes) performed on RGB videos, such as those captured by current AR headsets. This video must contain multiple views of the same scene but need not be captured by a stereoscopic camera. The following sections describe the primary components of the ARPOV interface, all of which are linked and interactive.

Augmented Reality User Interface


Illustration of the research project
AdaptiveCoPilot: Neuro-Adaptive Pre-Flight Guidance

To build an adaptive guidance system for aviation training, we developed AdaptiveCoPilot, a neuroadaptive, multimodal feedback system designed for a VR cockpit environment. This system dynamically adjusts task guidance based on pilots’ cognitive states, which are measured using fNIRS. By monitoring cognitive facets such as working memory, perception, and attention, AdaptiveCoPilot adapts the modality and content of feedback in real time, delivering visual, auditory, and text- based guidance tailored to the user’s workload. The design was informed by a formative study involving three pilots, which identified areas of high cognitive demand in aircraft operation and the challenges of selecting appropriate feedback modalities. These findings guided the development of neuroadaptive strategies to maintain optimal workload states and improve task performance during complex aviation procedures.

Illustration of the research project
Satori: Towards the Proactive Assistant

To build the adaptive UI, we propose a belief-desire-intention (BDI) user-modeling-based proactive assistance method called Satori, designed to dynamically adjust task guidance based on the user’s context, environment, and actions. This design draws upon insights from two formative studies aimed at identifying the challenges and opportunities in creating adaptive AR interfaces.

Illustration of the research project
ARTiST: AR Text Simplification for Task-Efficient UI

The AR interface must effectively support task performance. To achieve this, we focused on optimizing both text and graphic displays for each task step. To enhance text display, we implemented ARTiST, an automated text simplification system designed to produce shorter and more understandable instructions. This approach aims to improve task efficiency by ensuring that users can quickly grasp the necessary information without being overwhelmed by complex language.


Visualization Imaging and Data Analysis Center (VIDA Lab)
370 Jay Street 11th Floor, Brooklyn, NY 11201.