We lifted 3 functions out of this paper's own repositories and ran 2 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| homayoonfarrahi/cycle-time-study | canonical | 2 of 3 |
| Function | Status | Where it lives |
|---|---|---|
| create_log_gaussian | Ran | homayoonfarrahi/cycle-time-study/sac/utils.py code served (permissive licence) · get_code("2526cfd8d5bf5d06") |
| logsumexp | Ran | homayoonfarrahi/cycle-time-study/sac/utils.py code served (permissive licence) · get_code("09556d56c8b73050") |
| fc_body | Not yet run | homayoonfarrahi/cycle-time-study/rl/rl/nets/utils.py code served (permissive licence) · get_code("45cfa5d652612d16") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Continuous-time reinforcement learning tasks commonly use discrete steps of fixed cycle times for actions. As practitioners need to choose the action-cycle time for a given task, a significant concern is whether the hyper-parameters of the learning algorithm need to be re-tuned for each choice of the cycle time, which is prohibitive for real-world robotics. In this work, we investigate the widely-used baseline hyper-parameter values of two policy gradient algorithms -- PPO and SAC -- across different cycle times. Using a benchmark task where the baseline hyper-parameters of both algorithms were shown to work well, we reveal that when a cycle time different than the task default is chosen, PPO with baseline hyper-parameters fails to learn. Moreover, both PPO and SAC with their baseline hyper-parameters perform substantially worse than their tuned values for each cycle time. We propose novel approaches for setting these hyper-parameters based on the cycle time. In our experiments on simulated and real-world robotic tasks, the proposed approaches performed at least as well as the baseline hyper-parameters, with significantly better performance for most choices of the cycle time, and did not result in learning failure for any cycle time. Hyper-parameter tuning still remains a significant barrier for real-world robotics, as our approaches require some initial tuning on a new task, even though it is negligible compared to an extensive tuning for each cycle time. Our approach requires no additional tuning after the cycle time is changed for a given task and is a step toward avoiding extensive and costly hyper-parameter tuning for real-world policy optimization.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2305.05760")
get_code_for_paper("2305.05760")
have("2305.05760")
Connect an agent — have() is free.