Open Research Community

AI can answer. Can it learn from what happens next?

This community studies AI systems that act, change their environment, receive feedback, remember failure, and carry evidence into the next decision.

learnusecontributepartner
Diagram separating open public research artifacts from private production implementation.
ActWhat evidence must exist before an AI action proceeds?
ObserveWhat changes when an agent becomes part of the world it predicts?
LearnDoes a correction become durable, scoped behavior?
AuditCan another person inspect the rule, result, and limit?

Four Routes

Enter at the level that matches your question.

No affiliation or private access is required. Read first, run an artifact, improve something small, or bring a concrete reliability problem.

04

Partner

Bring a bounded evaluation problem where failure recurrence or action evidence has real cost.

Start a private inquiry

Research Projects

Four projects, each built around a concrete question.

Open the paper, protocol, schema, example, or reference implementation that matches the question you want to investigate.

Versioned Releases

Versioned artifacts with commands, manifests, and explicit limits.

Each release fixes the exact files, version, and verification commands needed to cite or run it. None is a production safety certificate or the complete product.

Current Research Conversation

Two KDD 2026 workshop papers. Two ways an agent can fail after it becomes consequential.

ReflexBench evaluates whether agent reasoning survives deeper observer-participant feedback. Across 20 scenarios, six domains, four observer-depth levels, and nine public models, all tested models degraded at deeper levels.

Cognitive Immunity evaluates repeated severe-threshold failure after correction. Under the documented score-level protocol, the pooled repeat failure rate moved from 0.764 for No Memory to 0.650 for Cognitive Immunity; the paired seed-task delta was -0.106 with a 95% bootstrap interval of [-0.211, -0.003].

Boundary: both are accepted and presented non-archival KDD 2026 workshop papers. They are not KDD main-conference proceedings papers, production safety certification, or proof of general reliability.

Mian Zhang with the ReflexBench and Cognitive Immunity posters at KDD 2026 in Jeju.

First Contribution

Participation can start small and become rigorous.

1Read or run one public artifact.
2Ask one concrete question or describe one use case.
3Repair one document, link, fixture, or example.
4Publish a baseline, experiment, or extension.
5Maintain a bounded component or research thread.

Open-Core Boundary

Open enough to inspect and extend. Bounded enough to protect people and product.

Public

Published concepts and formulas, protocols, schemas, validators, small fixtures, minimal references, safe demos, documentation, and negative results.

Release review

Large datasets, detailed traces, near-production references, security-relevant examples, and components whose combination may reveal implementation advantage.

Private

Production orchestration, exact weights and thresholds, private prompts and data, customer systems, tuning history, deployment automation, and unreleased research.

Read the full release boundary

Questions

What this community is, and what it is not.

Who can contribute?

Researchers, engineers, students, operators, and readers. Useful work includes questions, use cases, documentation, safe fixtures, baselines, experiments, and code.

Must every contribution reproduce an experiment or dispute a conclusion?

No. Contributions can clarify a concept, connect related materials, improve usability, add an example, or describe a real problem. Counterexamples and conclusion checks are two options among many.

Does open research mean the complete product is open source?

No. Public protocols, interfaces, validators, and minimal references are separated from private production implementation, operational parameters, customer material, and proprietary learning loops.

Does a paper or benchmark certify production safety?

No. Each result is bounded by its stated protocol, data, version, and limitations.

Community Standard

Bring a real question. Leave behind a public artifact another person can use.

The artifact may be a clearer explanation, a safe fixture, a comparison, a negative result, a corrected boundary, or a working extension.