Learn
Read plain-language concepts, papers, formulas, results, and limitations.
Start with the research mapOpen Research Community
This community studies AI systems that act, change their environment, receive feedback, remember failure, and carry evidence into the next decision.
Four Routes
No affiliation or private access is required. Read first, run an artifact, improve something small, or bring a concrete reliability problem.
Read plain-language concepts, papers, formulas, results, and limitations.
Start with the research mapInspect benchmarks, datasets, schemas, validators, demos, and minimal references.
Open Hugging Face artifactsAsk a question, add a safe example, improve a document, or publish an experiment.
Choose a discussion routeBring a bounded evaluation problem where failure recurrence or action evidence has real cost.
Start a private inquiryResearch Projects
Open the paper, protocol, schema, example, or reference implementation that matches the question you want to investigate.
Versioned Releases
Each release fixes the exact files, version, and verification commands needed to cite or run it. None is a production safety certificate or the complete product.
Current Research Conversation
ReflexBench evaluates whether agent reasoning survives deeper observer-participant feedback. Across 20 scenarios, six domains, four observer-depth levels, and nine public models, all tested models degraded at deeper levels.
Cognitive Immunity evaluates repeated severe-threshold failure after correction. Under the documented score-level protocol, the pooled repeat failure rate moved from 0.764 for No Memory to 0.650 for Cognitive Immunity; the paired seed-task delta was -0.106 with a 95% bootstrap interval of [-0.211, -0.003].
Boundary: both are accepted and presented non-archival KDD 2026 workshop papers. They are not KDD main-conference proceedings papers, production safety certification, or proof of general reliability.

First Contribution
Open-Core Boundary
Published concepts and formulas, protocols, schemas, validators, small fixtures, minimal references, safe demos, documentation, and negative results.
Large datasets, detailed traces, near-production references, security-relevant examples, and components whose combination may reveal implementation advantage.
Production orchestration, exact weights and thresholds, private prompts and data, customer systems, tuning history, deployment automation, and unreleased research.
Questions
Researchers, engineers, students, operators, and readers. Useful work includes questions, use cases, documentation, safe fixtures, baselines, experiments, and code.
No. Contributions can clarify a concept, connect related materials, improve usability, add an example, or describe a real problem. Counterexamples and conclusion checks are two options among many.
No. Public protocols, interfaces, validators, and minimal references are separated from private production implementation, operational parameters, customer material, and proprietary learning loops.
No. Each result is bounded by its stated protocol, data, version, and limitations.
Community Standard
Bring a real question. Leave behind a public artifact another person can use.
The artifact may be a clearer explanation, a safe fixture, a comparison, a negative result, a corrected boundary, or a working extension.