In this paper, we propose a model that utilizes input images from two forward-facing cameras, along with vehicle velocity and traffic light status to predict the future waypoints of the vehicle. Trained on expert demonstrations, the model learns to predict future waypoints in the vehicle frame of reference without access to BEV ground truths.