GSM8K

Explore our entire collection of insights, tutorials, and industry news.

  • AI Tutorials

    Karpathy nanochat GRPO RL Loop Deep Dive

    An in-depth technical analysis of Andre Karpathy's 300-line GRPO implementation in nanochat, exploring how simplified reinforcement learning can solve complex reasoning tasks like GSM8K.
    Read more