AAAI Conference (2023)
Lajanugen Logeswaran, Wilka Carvalho (University of Michigan), Honglak Lee
Abstract
The ability to combine learned knowledge and skills to solve novel tasks is a key aspect of generalization in humans that allows us to understand and perform tasks described by novel language utterances. While progress has been made in su- pervised learning settings, no work has yet studied compo- sitional generalization of a reinforcement learning agent fol- lowing natural language instructions in an embodied environ- ment. We develop a set of tasks in a photo-realistic simu- lated kitchen environment that allow us to study the degree to which a behavioral policy captures the systematicity in lan- guage by studying its zero-shot generalization performance on held out natural language instructions. We show that our agent which leverages a novel additive action-value decom- position in tandem with attention-based subgoal prediction is able to exploit composition in text instructions to generalize to unseen tasks.