{"author":"haibinlin","children":[{"author":"haibinlin","children":[],"created_at":"2025-02-06T18:58:50.000Z","created_at_i":1738868330,"id":42965362,"options":[],"parent_id":42965361,"points":null,"story_id":42965361,"text":"The weekly trending python project on github (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;trending&#x2F;python?since=weekly\">https:&#x2F;&#x2F;github.com&#x2F;trending&#x2F;python?since=weekly</a>) is *verl* (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;volcengine&#x2F;verl\">https:&#x2F;&#x2F;github.com&#x2F;volcengine&#x2F;verl</a>)<p>This framework is used by Bytedance to train RL reasoning models that achieves openAI O1 level performance on olympics math problems. Code is open sourced.<p>Now many people are building on top of it to reproduce deepseek R1:\n- <a href=\"https:&#x2F;&#x2F;github.com&#x2F;Jiayi-Pan&#x2F;TinyZero\">https:&#x2F;&#x2F;github.com&#x2F;Jiayi-Pan&#x2F;TinyZero</a> (8.6k stars): reproducing ds-r1 zero on countdown tasks for just $30, with a model with 3 billion parameters\n- <a href=\"https:&#x2F;&#x2F;github.com&#x2F;ZihanWang314&#x2F;ragen\">https:&#x2F;&#x2F;github.com&#x2F;ZihanWang314&#x2F;ragen</a> (700 stars): ex-deepseek researcher use verl to build open source LLM + RL + agent model training framework<p>Reinforcement learning is back as the hottest topic now. Why was it not popular in previous years?","title":null,"type":"comment","url":null}],"created_at":"2025-02-06T18:58:50.000Z","created_at_i":1738868330,"id":42965361,"options":[],"parent_id":null,"points":4,"story_id":42965361,"text":null,"title":"OSS reinforcement learning lib by ByteDance is used to reproduce DeepSeek R1","type":"story","url":"https://github.com/volcengine/verl"}
