Are benchmark results available?
No. The current report is a draft protocol. It deliberately contains no public performance numbers.
Evidence before claims
This page publishes the measurement plan before the results. Draft values are null and cannot be rendered as performance claims.
The benchmark has a real-world layer using each product's normal configuration and a controlled layer using the same model and corpus for Kosmico and a conventional multi-tool stack.
The real-world layer compares the experience a researcher would normally receive from Kosmico, ChatGPT Deep Research, NotebookLM, and Elicit. The controlled layer isolates the workspace by holding the model and corpus constant against a conventional combination of model, academic search, Zotero, and Overleaf.
Four workflows run across five disciplinary topics with three repetitions per condition. Two blinded expert evaluators score the artifacts, disagreements are adjudicated, and the raw runs, rubric, environment manifest, and sample sizes are retained.
A result can be published only when every metric has a finite value, the raw-run location and publication date exist, and evidence review has approved the claim. Until then, the public report exposes the protocol and null results only.
Benchmark status
The protocol is public, but no results are published. Every result field remains null until evidence review and publication of the raw runs.
4 workflows, 5 topics, 3 repetitions per condition, 2 blinded expert evaluators.
37
Workflow, autonomy, context, retrieval, evidence, stopping, reliability, cost, and LLM visibility measures. Public values: none while draft.
No. The current report is a draft protocol. It deliberately contains no public performance numbers.
It helps separate the effect of the workspace and research pipeline from differences between underlying models or corpora.
Publication requires a raw-run URL, the rubric, environment information, sample sizes, and the stated expert-evaluation process.