Understand the claim.
Identify what supports a conclusion: a fair comparison, statistical evidence, unseen test data, or an experiment that covers the full workload.
“Our method reliably improves accuracy across independent runs.”
Bring the report, code, and experiments into one review. We’re building an AI assistant to help professors connect a project’s conclusions to the evidence—and find the follow-up questions that make the work stronger.
In development · Contact us for testing & feedback
4.2 Results and evaluation
“Retrieval improves accuracy from 67.8% to 74.3% under identical evaluation.”
A clearer way to review
In a large project, the evidence is spread across data pipelines, code paths, evaluation scripts, and saved runs. Connecting those pieces can make the next review more focused and productive.
Identify what supports a conclusion: a fair comparison, statistical evidence, unseen test data, or an experiment that covers the full workload.
“Our method reliably improves accuracy across independent runs.”
Trace the relevant files, data boundaries, configurations, and runtime evidence to see which conditions actually produced the result.
claim → run → config → code path
Turn the evidence into a focused review question, with references and a comparison, trace, or experiment that could strengthen the conclusion.
“Does the interval for the paired difference support a positive gain?”
Explore the review approach
Explore six project review examples. See how a review can connect a conclusion to its evidence and suggest a practical next step.
Evidence to inspect.
Questions to move the work forward.
“Retrieval improves accuracy from 67.8% to 74.3% under identical evaluation.”
# runs/baseline.yaml
questions: qa_test_v1.jsonl # 500 items
harness: qa-eval@1.3
scorer: raw_exact_match
accuracy: 0.678
# runs/retrieval.yaml
questions: qa_test_filtered.jsonl # 420 items
harness: qa-eval@1.4
scorer: normalized_exact_match
accuracy: 0.743The scores match the report, but the question set and scoring rule changed. A matched evaluation would clarify how much of the difference comes from the system itself.
Re-evaluate both systems on the same question IDs with one harness and scoring rule. Does the reported gain remain?
Built around the reviewer
Our goal is to help supervisors spend more time on the reasoning that matters. Each review prompt should connect the evidence, explain its significance, and offer a useful direction for the next discussion.
Test the approach with usThe intended output links each review question to a report passage and the relevant code, data, or experiment.
Show what supports a conclusion and where more context or another experiment would help.
Give supervisors focused questions to explore with their teams and a practical path toward a stronger final report.
ReportProof is in development, and we’re inviting professors and research supervisors to help shape it. Get in touch to explore early testing, share your review workflow, and tell us what would be useful.
Contact us to testA few practical questions
Email contact@reportproof.online to explore early testing. Tell us about the projects you review and what you’d like help checking. We’re developing ReportProof with early testing coordinated directly, and your feedback can help shape the experience.
We’re starting with ML and AI projects that combine a written report with code, data pipelines, run configurations, and saved experiment outputs. The focus is claims whose support is spread across several parts of the project.
The focus is whether the evidence supports the conclusion. Examples include statistical uncertainty that conflicts with a confident claim, a reported test setup that differs from the executed code, changed evaluation conditions, test answers entering a retrieval index, excluded failures, and experiments that change several variables at once.
Use it to start a closer look. Follow the references, consider the project’s context, and discuss the suggested check with the team. The goal is to help supervisors assess the evidence and decide which follow-ups will strengthen the work.
For early testing, we’ll agree on suitable project material, permissions, and data handling together. You can start by describing your workflow over email. The interactive examples on this website use fictional project content.