ANALYSIS PSPA-Bench: New Benchmark Exposes Personalization Gap in Smartphone GUI Agents 3/10 4 min read 6 months ago
ANALYSIS Frontier Models Hit 19% Meltdown Rate in Long-Horizon LLM Agent Study 4/10 4 min read 6 months ago
ANALYSIS RIDE Study: Routing Meta Prompts Densify LLM Layers, Not Sparsify 3/10 4 min read 6 months ago
ANALYSIS Webscraper Framework Uses MLLMs to Extract Data From Dynamic Sites 4/10 4 min read 6 months ago
ANALYSIS Los Alamos Researchers Use LLMs to Build Biodefense Countermeasure Databases 4/10 4 min read 6 months ago
ANALYSIS SciVisAgentBench: 108 Cases for Testing Scientific Visualization Agents 3/10 4 min read 6 months ago
ANALYSIS GISTBench Tests LLM User Understanding in Recommendation Systems 3/10 4 min read 6 months ago
ANALYSIS PAR²-RAG Framework Beats IRCoT by 23.5% on Multi-Hop Question Answering 3/10 4 min read 6 months ago
ANALYSIS Self-Organizing LLM Agents Outperform Designed Structures by 14%, Study Finds 5/10 4 min read 6 months ago
ANALYSIS Mimosa Multi-Agent Framework Achieves 43.1% on ScienceAgentBench 4/10 4 min read 6 months ago