[Cognitive Revolution] The Tiny Model Revolution with Ronen Eldan and Yuanzhi Li of Microsoft Research

Latent Space: The AI Engineer Podcast ·

Thanks to the over 1m people that have checked out the Rise of the AI Engineer. It’s a long July 4 weekend in the US, and we’re celebrating with a podcast feed swap! We’ve been big fans of Nathan Labenz and Erik Torenberg’s work at the Cognitive Revolution podcast (https://www.cognitiverevolution.ai/) for a while, which started around the same time as we did and has done an incredible job of hosting discussions with top researchers and thinkers in the field, with a wide range of topics across computer vision (https://www.cognitiverevolution.ai/e6-the-computer-vision-revolution-with-junnan-li-and-dongxu-li-of-blip-and-blip2/) (a special focus thanks to Nathan’s work at Waymark), GPT-4 (https://www.cognitiverevolution.ai/e11-openais-gpt-4-discussion-with-nathan-labenz-and-erik-torenberg/) (with exceptional insight due to Nathan’s time on the GPT-4 “red team (https://twitter.com/labenz/status/1660765244363350018)”), healthcare/medicine/biotech (Harvard Medical School (https://www.cognitiverevolution.ai/e24-the-ai-revolution-in-medicine-with-dr-isaac-kohane-of-harvard-medical-school/), Med-PaLM (https://www.cognitiverevolution.ai/e27-googles-med-palm-and-med-palm2-with-vivek-natarajan/), Tanishq Abraham (https://www.cognitiverevolution.ai/e38-the-virtual-biopsy-revolution-with-dr-tanishq-mathew-abraham-part-2-of-2/), Neal Khosla (https://www.cognitiverevolution.ai/e25-revolutionizing-patient-care-with-neal-khosla-of-curai-health/)), investing and tech strategy (Sarah Guo, Elad Gil (https://www.cognitiverevolution.ai/e21-vc-insights-on-investing-in-artificial-intelligence-with-sarah-guo-and-elad-gil-of-no-priors-podcast/), Emad Mostaque, Sam Lessin (https://www.cognitiverevolution.ai/e35-stability-ais-emad-mostaque-and-slow-ventures-sam-lessin-discuss-investing-in-ai/)), safety (https://www.cognitiverevolution.ai/e26-bonus-episode-connor-leahy-on-agi-gpt-4-and-cognitive-emulation-w-fli-podcast/) and policy (https://www.cognitiverevolution.ai/e16-pausing-the-ai-revolution-with-technologist-jaan-tallinn/), curators (https://www.cognitiverevolution.ai/e23-scouting-the-ai-revolution-with-robert-scoble-and-bens-bites-creator-ben-tossell/) and influencers (https://www.cognitiverevolution.ai/e22-helping-businesses-use-ai-with-rachel-woods-of-the-ai-exchange/) and exceptional AI founders (Josh Browder (https://www.cognitiverevolution.ai/e34-the-consumer-rights-revolution-with-joshua-browder-of-donotpay/), Eugenia Kuyda (https://www.cognitiverevolution.ai/e3-the-empathy-revolution-with-eugenia-kuyda-of-replika/), Flo Crivello (https://www.cognitiverevolution.ai/e9-the-ai-assistant-revolution-with-flo-crivello-of-lindyai/), Suhail Doshi (https://www.cognitiverevolution.ai/e1-the-pixel-revolution-with-playgroundais-suhail-doshi/), Jungwon Byun (https://www.cognitiverevolution.ai/e14-the-reasoning-revolution-with-oughts-jungwon-byun-and-andreas-stuhlmuller/), Raza Habib (https://www.cognitiverevolution.ai/e20-the-great-implementation-with-raza-habib-ceo-of-humanloop/), Mahmoud Felfel (https://www.cognitiverevolution.ai/e10-the-ai-voice-revolution-with-mahmoud-felfel-of-playht/), Andrew Feldman (https://www.cognitiverevolution.ai/e29-the-ai-chip-revolution-with-andrew-feldman-of-cerebras/), Matt Welsh (https://www.cognitiverevolution.ai/e19-the-ai-agent-revolution-with-matt-welsh-of-fixieai/), Anton Troynikov (https://www.cognitiverevolution.ai/e5-the-embedding-revolution-anton-troynikov-on-chroma-stable-attribution-and-future-of-ai/), Aravind Srinivas (https://www.cognitiverevolution.ai/e7-in-search-of-truth-with-aravind-srinivas-of-perplexity-ai/)). If Latent Space is for AI Engineers (https://www.latent.space/p/ai-engineer), then Cognitive Revolution covers the much broader field of AI in tech, business and society at large, with a longer runtime to go deep on research papers like TinyStories. We hope you love this episode as much as we do, and check out CogRev wherever fine podcasts are sold! Subscribe to the Cognitive Revolution on:

  • Website
  • Apple Podcasts
  • Spotify
  • Youtube Good Data is All You Need The work of Ronen and Yuanzhi echoes a broader theme emerging in the midgame of 2023:
  • Falcon-40B (trained on 1T tokens) outperformed (https://twitter.com/johannes_hage/status/1662119769045188610) LLaMA-65B (trained on 1.4T tokens), primarily due to the RefinedWeb Dataset (https://arxiv.org/abs/2306.01116) that runs CommonCrawl through extensive preprocessing and cleaning in their MacroData Refinement pipeline.
  • UC Berkeley LMSYS’s Vicuna-13B is near GPT-3.5/Bard quality (https://lmsys.org/blog/2023-03-30-vicuna/) at a tenth of their size, thanks to fine-tuning from 70k user-highlighted ChatGPT conversations (indicating some amount of quality).
  • Replit’s finetuned 2.7B model outperforms the 12B OpenAI Codex model (https://www.latent.space/p/reza-shabani#details) based on HumanEval, thanks to high quality data from Replit users The path to smaller models leans on better data (and tokenization!), whether from cleaning, from user feedback, or from synthetic data generation (https://twitter.com/karpathy/status/1671587087542530049?s=20), i.e. finetuning high quality on outputs from larger models. TinyStories and Phi-1 are the strongest new entries in that line of work, and we hope you’ll pick through the show notes to read up further.

Show Notes

This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe (https://www.latent.space/subscribe?utm_medium=podcast&utm_campaign=CTA_2)