Evidence before claims

Research-workflow benchmark

This page publishes the measurement plan before the results. Draft values are null and cannot be rendered as performance claims.

The benchmark has a real-world layer using each product's normal configuration and a controlled layer using the same model and corpus for Kosmico and a conventional multi-tool stack.

Two layers answer different questions

The real-world layer compares the experience a researcher would normally receive from Kosmico, ChatGPT Deep Research, NotebookLM, and Elicit. The controlled layer isolates the workspace by holding the model and corpus constant against a conventional combination of model, academic search, Zotero, and Overleaf.

The work must be reproducible

Four workflows run across five disciplinary topics with three repetitions per condition. Two blinded expert evaluators score the artifacts, disagreements are adjudicated, and the raw runs, rubric, environment manifest, and sample sizes are retained.

No number before evidence approval

A result can be published only when every metric has a finite value, the raw-run location and publication date exist, and evidence review has approved the claim. Until then, the public report exposes the protocol and null results only.

Benchmark status

Draft protocol

draft

The protocol is public, but no results are published. Every result field remains null until evidence review and publication of the raw runs.

Protocol

4 workflows, 5 topics, 3 repetitions per condition, 2 blinded expert evaluators.

Conditions

  • Kosmico default configuration (real-world)
  • ChatGPT Deep Research (real-world)
  • NotebookLM (real-world)
  • Elicit (real-world)
  • Kosmico controlled condition (controlled)
  • Model + academic search + Zotero + Overleaf (controlled)

Metrics

37

Workflow, autonomy, context, retrieval, evidence, stopping, reliability, cost, and LLM visibility measures. Public values: none while draft.

Retained evidence

  • raw runs
  • evaluation rubric
  • environment manifest
  • sample sizes

Direct answers

Are benchmark results available?

No. The current report is a draft protocol. It deliberately contains no public performance numbers.

Why include a controlled comparison?

It helps separate the effect of the workspace and research pipeline from differences between underlying models or corpora.

Will raw runs be public?

Publication requires a raw-run URL, the rubric, environment information, sample sizes, and the stated expert-evaluation process.