All tags
Person: "wtgowers"
not much happened today
raev2 gated-deltanet-2 kda mamba-3 dclm nvidia openai nous-research representation-learning tokenization linear-attention long-context mechanistic-interpretability math data-filtering agent-infrastructure language-modeling commonsense-reasoning 1jaskiratsingh recatm sainingxie ahatamiz1 rasbt nousresearch tatsu_hashimoto goodfireai markchen90 wtgowers memecrashes cloneofsimo lvwerra
RAEv2 advances representation-first tokenization with >10x faster convergence and improved generation, tested on text-to-image and world models. NVIDIA's Gated DeltaNet-2 innovates linear attention with channel-wise gates, outperforming KDA and Mamba-3 at 1.3B parameters on language modeling and reasoning tasks. Studies on subword tokenization reveal only some benefits at scale, while data filtering research suggests that with enough compute, no filtering may be optimal at around 1e30 FLOPs. Mechanistic interpretability updates propose clustering features by joint firing patterns for better geometry understanding. OpenAI's AI-assisted breakthrough on an Erdős unit-distance math problem sparks debate on AI's role in mathematical research. Harnesses remain key for capability improvements in agent infrastructure.
not much happened today
command-a+ claude-3.7-sonnet openai cohere reinforcement-learning reasoning multimodality model-architecture model-optimization model-releases benchmarking long-context model-efficiency transformers wtgowers hongxunwu aidangomez nickfrosst clementdelangue eliebakouch rasbt sama
OpenAI achieved a major math breakthrough by disproving a long-standing Erdős unit distance problem using a general-purpose reasoning model, marking a milestone in AI-driven formal science and long-horizon reasoning. The result was validated by prominent mathematicians like Timothy Gowers and OpenAI researcher Hongxun Wu, highlighting the model's advanced reasoning capabilities beyond prior AI math achievements. Meanwhile, Cohere released Command A+ as an open-source Apache 2.0 licensed model, featuring a 218B MoE / 25B active multimodal architecture supporting 48 languages and optimized for low hardware requirements, runnable on as little as 2× H100 GPUs. Benchmarks place Command A+ near Claude 4.5 Haiku in intelligence with strong non-hallucination but weaker scientific reasoning and coding. The architecture includes novel elements like a parallel transformer block, shared experts, and LayerNorm over RMSNorm.