Offline-to-online reinforcement learning pipelines may not need pretrained Q-functions: a new Stanford preprint by Chelsea ...