AGENTENCODE / RESEARCH & ORCHESTRATION

Methodology: Verifiable Multi-Agent Research

A transparent evidence contract for studies, benchmark scope, provenance and limitations.

Editorial: AgentenCodeReviewed: 9 October 2026Primary sources & external studies
External research is not presented as our own measurement. Unknown remains unknown.

How does AgentenCode verify multi-agent claims?

The Multi-Agent Intelligence methodology separates five auditable layers: original source, exact claim, documented or measured scope, review status and subsequent change history. Vendor documentation may establish that a framework describes a workflow pattern. A peer-reviewed paper may establish that a result was reported under specified experimental conditions. Neither makes that finding an independent AgentenCode measurement. The first public research registry records the evidence model and its reviewed entries.

Provenance categories and what they mean

evidence_origin Meaning Defensible statement Common limitation
peer_reviewed_external Refereed publication A study reports a result under its setup No automatic deployment-wide generalization
independent_external Third-party independent test External measurement with documented conditions Validate reproducibility and interests
vendor_reported Supplier statement Provider asserts a scoped property or figure Not independently demonstrated
project_documentation Framework or standards documentation A method, API or workflow is described Not an empirical performance comparison
agentencode_measured Repeatable in-house AgentenCode test Only from actually archived test executions No such measurements in version 1

Verification status is a separate dimension: reviewed, pending, conflicting or insufficient_evidence. A peer-reviewed paper can be a trusted publication while still providing insufficient evidence for an unrelated product-level statement. “Reviewed” means the exact cited claim was inspected, not that the architecture received a general safety certification.

Capabilities follow the values yes, no, unknown and not_applicable. Unknown is never equivalent to no or yes. A negative product property requires an explicit source for the relevant scope. Region, plan, integration type and runtime can all change the interpretation of privacy or security claims.

Research data contract: five entity families

  1. Source: persistent source_id, original URL/DOI, publisher, source kind, publication and retrieval dates, rights status and original-content hash where genuinely available.
  2. BenchmarkStudy: study_id, research question, tasks, dataset revisions, design, comparison groups and limitations.
  3. EvaluationResult: result_id, linked study, specific model/topology/setup, metric and units, value, uncertainty, provenance and status. There are no fabricated numeric results where extraction was not verified.
  4. ArchitecturePattern: explicit control flow, message/state arrangements, common failure modes and evaluation questions.
  5. EvidenceLink: one claim tied to the exact source, passage, review decision, scope and any subsequent version history.

The machine-readable research dataset contains a small reviewed initial set. It is not a wholesale import of third-party research. A separate graph projection contains explicit supported edges only; there is no reasoning from “framework X implements Y” to “every agent using the platform supports Y”.

Source-to-claim review workflow

1. Select authoritative sources. For academic findings prefer the official publication and DOI, for implemented functions the maintaining project's documentation, and for protocols a versioned specification. Code repositories help reproduction; they are not substitutes for peer-reviewed experimental designs.

2. Record the specific passage. Store page, section or stable fragment where practical. MultiAgentBench statements about graph topology and cognitive planning concern the ACL 2025 experiment. They are not evidence about every AI ecosystem catalogued by AgentenCode.

3. Type the proposition. Is it externally measured, technically documented, interpreted or pending investigation? A verification=pending observation must not appear as a fully established result.

4. Preserve the full scope. Task family, tested agent count, topology, model/version, dataset, tool permissions and known runtime factors constrain the meaning of an evaluation. Missing details stay unknown.

5. Expose conflicting evidence. When primary sources disagree, preserve both. Do not overwrite historical studies just because a specification or model was revised. A redirect or changed documentation URL is a monitoring signal first, not a product release event.

6. Review before public comparison. Metric units, comparable baseline, statistical sample, data rights and dependency versions must pass explicit gates. Without those gates numeric entries are excluded from direct side-by-side comparison.

Methodological comparability and uncertainty

Two scores can be placed in a direct quantitative comparison only when task, dataset version, metric definition, model/system configuration, time and allowed compute/tool budget are sufficiently compatible. Different success criteria, tools or prompts are potential confounders. Number of trials, variance, significance method and excluded cases also matter. Unreported uncertainty is unknown, never zero by default.

The MultiAgentBench paper and AgentBench implementation address different evaluation constructs. AgentenCode deliberately does not invent a unified score. This is a scientific quality constraint, not a missing chart.

The ACL Anthology describes CC BY 4.0 as its default licensing for materials from 2016 onward; embedded components can carry separate rights. The MARBLE repository identifies an MIT license for its code. Neither is a blanket license to redistribute unrelated third-party datasets, exported result tables or figures. We favor original analysis, limited quotation, attribution and authoritative links. Before any automated import, access terms, API limits, revision stability and component-level licenses need separate review.

Sources contain actual review timestamps. Where the bytes of an external original document were not obtained and hashed, the original-content hash is explicitly null with an explanation: a local metadata digest is not misrepresented as the external source hash.

AgentenGraph, AgentenWache and temporal evidence

The new research graph is released as a separate explicitly sourced projection. For example the documented Microsoft Agent Framework scope may connect to a specific orchestration pattern. That does not show that every Copilot product implements the same capability. The original AgentenGraph relations remain unchanged.

The source registry includes candidates for future AgentenWache monitoring of DOI records, documentation, specifications and selected project releases. Candidates are not falsely described as running background jobs. Operational monitoring needs source-state storage, throttling, meaningful-change detection and review gates. Existing daily News, Evidence, Changelog and agent checks remain untouched. Monitoring signals must be validated before becoming public facts.

What an actual AgentenCode experiment would require

An in-house study would need a registered protocol covering exact tasks, permitted data reuse, single-agent baseline, pinned model and tool versions, equal budgets, seeds, execution environment, safety limits, archived traces, costs and statistical processing. Only observed and reproducible executions could receive the label agentencode_measured. There are no in-house benchmark results in this release. The benchmark page remains a curated external-research resource.

Corrections and update frequency

Meaningful corrections update the affected text and that URL's truthful lastmod. A routine daily check is not a fresh measurement. The existing agent history is not padded with product events just because the editorial research area is launched. Refer to the general methodology and public history for the rest of the platform.

Primary sources and provenance

Each citation supports a scoped proposition. Third-party results have not been independently measured by AgentenCode.

  1. Zhu et al. · MultiAgentBench, ACL 2025
  2. MARBLE · research implementation
  3. AgentBench · reference repository
  4. Microsoft Agent Framework · Orchestrations
  5. A2A · protocol specification

Research dataset: JSON · explicit graph edges: JSON

Frequently asked questions

How can I identify vendor-reported information?

The evidence_origin field distinguishes vendor_reported and project_documentation from peer-reviewed or independent external work.

Does Unknown mean a feature is missing?

No. Unknown means the reviewed evidence does not establish the proposition within the stated scope.

Are the assessments independent?

Provenance and review status are recorded, but an editorial source review is not equivalent to an independent AgentenCode experiment.

When is a study record updated?

After a confirmed version change, relevant correction, method revision or reviewed conflict—not merely because a calendar day changed.

Where are the original measurements?

At the cited primary sources. Version 1 does not automatically redistribute full third-party result tables.