← Mario Seddik

Evaluating AI for earnings-call preparation

I built agents that predict the questions analysts will ask on an earnings call, then evaluate those predictions against the actual transcript. Across 18 held-out calls, the predictions achieved 74% semantic question coverage.

Recreated interface with representative demo data, not client material.

The problem

Investor-relations teams prepare executives for the question-and-answer portion of earnings calls. Past transcripts provide a way to test whether AI improves the breadth of preparation: generate predictions from evidence available beforehand, then compare them with what analysts actually asked.

How it works

A coverage panel of agents reviews the evidence and predicts likely topics. A second group models the interests of analysts who cover the company. A synthesizer and deterministic merge combine the predictions into a ranked slate. The web interface streams the results and lets a user compare them with the real call.

The predicted slates covered 74% of actual questions across 18 held-out 2025 calls. The result measures semantic topic coverage on that sample using model-assisted grading.

The AI

Inputs include management’s presentation, prior calls, and relevant peer discussions. The target call’s Q&A is held out. Grading uses model judges, partial-credit rubrics, and multiple votes to measure semantic coverage. This assesses whether predicted questions address the topics that arose; it is not exact wording or asker-identification accuracy.

If an agent fails, the app surfaces the error. It does not replace the failed run with a canned prediction. Separating the evidence pack, prediction step, and grading step also makes the process easier to inspect and repeat.

Stack

Python, headless Claude agents, multi-agent orchestration, NDJSON streaming, model-judge evaluation harness, single-page web UI.