POLICYGRID: Causal Discovery for Adaptive Policy Optimization in Embodied Agents (Student Abstract)

Taqiya Ehsan; Shuren Xia; Jorge Ortiz

doi:10.1609/aaai.v40i48.42214

POLICYGRID: Causal Discovery for Adaptive Policy Optimization in Embodied Agents (Student Abstract)

Authors

Taqiya Ehsan Rutgers University
Shuren Xia Rutgers University
Jorge Ortiz Rutgers University

DOI:

https://doi.org/10.1609/aaai.v40i48.42214

Abstract

Embodied agents must reason causally, as correlation-based models fail under intervention and distribution shift. This challenge arises in domains like robotics and cyber-physical systems, where agents balance efficiency and comfort under uncertainty. We introduce POLICYGRID, unifying causal discovery and control by treating each action as both decision and experiment. Leveraging constraint-based search, neural causal models, and language model priors with interventional validation, POLICYGRID yields adaptive, interpretable policies. Across synthetic, real-world, and live deployments, it achieves superior causal recovery (F1 = 0.89) and 2.8× better multi-objective performance than correlation-based baselines, demonstrating safe, generalizable decision-making.

AAAI-26 / IAAI-26 / EAAI-26 Proceedings Cover

Downloads

PDF
Poster

Published

2026-03-14

How to Cite

Ehsan, T., Xia, S., & Ortiz, J. (2026). POLICYGRID: Causal Discovery for Adaptive Policy Optimization in Embodied Agents (Student Abstract). Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 41203–41205. https://doi.org/10.1609/aaai.v40i48.42214

Download Citation

Issue

Vol. 40 No. 48: EAAI-26 AI for Education, Model AI Assignments, AAAI-26 Emerging Trends, Doctoral Consortium, Student Abstracts, Undergraduate Consortium and Demonstrations

Section

AAAI Student Abstract and Poster Program

POLICYGRID: Causal Discovery for Adaptive Policy Optimization in Embodied Agents (Student Abstract)

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information