TRPO is an AI model developed by University of California (UC) Berkeley (United States), first published in February 2015. It works in the games domain, on tasks such as atari.
Epoch AI has no training-compute estimate for this model. The model has 33,500 parameters.
Access: Unreleased. Its weights are not openly released. The reference paper has 8,305 citations. Epoch AI rates the confidence of this record as confident.