Reconstruction is a blind benchmark designed to test whether large language models (LLMs) can accurately recover the core research idea of a published paper...
When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding investigates how teams of AI coding agents interact while completing programming task...
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy introduces a framework for Vision-Language-Action (VLA) models t...
Quipu is an embeddable knowledge graph store designed to manage data written by software agents.
Policy Iteration with Human Feedback (PIHF) is a framework designed to improve the diagnostic accuracy of language models in rare-disease cases by using a hu...
Chronocooked is a reinforcement learning (RL) benchmark suite designed to evaluate how artificial agents develop an internal sense of time.
This paper investigates whether language model compliance detectors actually evaluate specific regulatory rules or if they simply react to the general conten...
Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment addresses the scarcity of recognition resources for the wor...
LAVA (Logic-Aware Validation and Augmentation) is a modular framework designed to automate the auditing of complex financial documents, such as tax forms, ba...
GRIP (Grounded Reasoning via Information-Restricted Premises) is a method designed to improve how language models use retrieved evidence in retrieval-augment...