Последние посты

Just links
21 сент., 23:56
Small-Scale Experiments: Are We There Yet? https://arxiv.org/abs/2608.11859arXiv.orgSmall-Scale Experiments: Are We There Yet?Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and...
Just links
20 сент., 11:03
Classification of Rational c=1 Vertex Operator Algebras and Vertex Operator Superalgebras https://arxiv.org/abs/2609.00122arXiv.orgClassification of Rational $c=1$ Vertex Operator Algebras and...For mathematicians: In this first in a series of two papers, we give a mathematically rigorous classification of (sufficiently nice) $c=1$ vertex operator algebras (VOAs) and vertex operator...
Just links
17 сент., 11:36
Choosing the Trainable Geometry for Agentic RL https://www.trajectory.ai/field-notes/choosing-the-trainable-geometry-for-agentic-rlTrajectoryResearch lab and product company building the platform for continual learning.1,380Открыть в Telegram
Just links
16 сент., 20:04
https://fixupx.com/qingyu_shi_/status/2100181567666901052FixupXQingyu (@qingyu_shi_)A list of tasks for which LLMs (mostly GPT-6 Pro) have found solutions significantly better and non-trivially different from the authors’ solutions: https://qoj.ac/blog/qingyu/blog/4412 The list is still being updated, and I'll mark all solutions I find particularly…
Just links
16 сент., 18:02
https://github.com/rust-lang/rust/pull/156216 If you wanted to do for loop in const in Rust now you can (in nightly)GitHubimplement const Iterator for Range by Randl · Pull Request #156216 · rust-lang/rustView all comments This allows the use of for i in i..n in const. This I believe is interesting enough to be justified under #155816 r? @oli-obk1,170Открыть в Telegram
Just links
16 сент., 15:25
Dream-RSI: Recursive Self-Improvement through Evolving Worlds https://arxiv.org/abs/2609.14858arXiv.orgDream-RSI: Recursive Self-Improvement through Evolving WorldsRecursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is...
Just links
16 сент., 14:27
https://fixupx.com/creus_roger/status/2100080651214971017FixupXRoger Creus Castanyer (@creus_roger)astra playing NetHack: - score: 48,978 - BALROG progress: 49.4% - max dungeon depth: 17 - max experience level: 14 important note: this is a cherry-picked episode (best of 10), and Astra has been able to modify its own game-playing harness no privileged…
Just links
16 сент., 04:17
Nature Is Our Learning Environment https://periodic.com/news/nature-is-our-learning-environmentPeriodicNature Is Our Learning Environment – Periodic LabsScientific midtraining and reinforcement learning on our lab data make Periodic Neon the Pareto-optimal model for challenging XRD analysis.
Just links
15 сент., 20:26
Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble https://arxiv.org/abs/2609.10883arXiv.orgStory Imprinting: AI Assistants Absorb Traits from Human...Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's...1,040Открыть в Telegram
Just links
15 сент., 19:59
CheatBench: Measuring Reward Gaming in AI Agents https://www.cheatbench.ai/1,020Открыть в Telegram
Just links
15 сент., 19:12
https://fixupx.com/jsuarez/status/2099512196304925161Beating AutoAscend’s NetHack Challenge Score with RLNetHack is a game of long horizons, partial observability, and brutal interactions that can end a run with little warning. This is rough if you're learning to play, but makes for an unbelievably
Just links
15 сент., 18:59
https://fixupx.com/jsuarez/status/2099512196304925161The 5 Coolest RL Results in PufferLib 5.05. Multitask Drone Policies One single policy is able to both race and arrange into several different formations. Brought to you by our intern, Fin. Or should I say fintern... very appropriate for
Just links
15 сент., 10:28
An algorithm to generate two-dimensional critical lattice models using competing anyon condensation https://www.nature.com/articles/s41567-026-03438-6NatureAn algorithm to generate two-dimensional critical lattice models using competing anyon condensationNature Physics - Conformal field theories have a central role in statistical and high-energy theoretical physics. An algorithm to systematically generate critical lattice models described by...1,070Открыть в Telegram
Just links
15 сент., 10:21
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks https://arxiv.org/abs/2609.13009arXiv.orgHow Good Are Frontier Models at Physics? Expert Re-Grading Reveals...Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced...1,030Открыть в Telegram
Just links
14 сент., 18:29
переслано из @sqtech

Just links
13 сент., 20:25
Convergence Accelerators, or Why Most Foundation Model Findings from Academia Don’t Survive https://mlechner.substack.com/p/convergence-accelerators-or-why-mostSubstackConvergence Accelerators, or Why Most Foundation Model Findings from Academia Don’t SurviveLet me get one thing out there first.
Just links
11 сент., 18:26
https://neocognition.io/blog/apprentice-bench/NeoCognitionApprenticeBenchThe first end-to-end benchmark of computer use, continual learning, and long-horizon agentic capabilities, set in a real job.1,450Открыть в Telegram
Just links
11 сент., 18:11
https://chessbench-ai.github.io/chessbench-ai.github.ioChessBench — AI Chess LeaderboardCompare AI models by ChessBench rating. Explore benchmark results, methodology, and complete model-vs-model chess replays.1,450Открыть в Telegram
Just links
11 сент., 04:59
Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain https://arxiv.org/abs/2604.08407arXiv.orgYour Agent Is Mine: Measuring Malicious Intermediary Attacks on...Large language model (LLM) agents increasingly rely on third-party API routers to dispatch tool-calling requests across multiple upstream providers. These routers operate as application-layer...
Just links
11 сент., 04:46
https://cognition.com/blog/factoring-rsa-260CognitionFactoring RSA-260How Devin and a Cognition researcher built the world’s highest-performance GPU lattice siever, to make factoring numbers 10x cheaper than the previous…
