Media Summary: When a new AI model drops, it's judged based on a static benchmark grid that doesn't account for how long the model is allowed ... The lecture also covers a paper on optimally scaling Pre-training scaling laws are hitting a wall. Simply adding billions of parameters and petabytes of text data no longer yields linear ...
Really Big Test Time Compute - Detailed Analysis & Overview
When a new AI model drops, it's judged based on a static benchmark grid that doesn't account for how long the model is allowed ... The lecture also covers a paper on optimally scaling Pre-training scaling laws are hitting a wall. Simply adding billions of parameters and petabytes of text data no longer yields linear ... Build your voice AI agent today: Join My Newsletter for Regular AI Updates ... Is "thinking longer" better than simply training larger models? We analyze the first Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: ...
... a PhD student at UC Berkeley, shares two of his recent works: Scaling LLM Compute Optimal Test Time Scaling for Large Language Models Xuan-May Le, Minh-Tuan Tran, Ling Luo, Uwe Aickelin, Dinh Phung, Trung Le.