I was the primary builder of ALM Frontier: a platform that brings insurance data, financial models, and analytical applications into a shared working environment. I connected the workflows across them and built an AI layer that helps teams investigate results, test scenarios, and turn complex analysis into decisions.
The vision
Life and annuity teams work across policy administration, actuarial modeling, investments, finance, and reporting. Each function depends on the others, yet their data and tools often sit apart. I built ALM Frontier around a connected workflow: prepare the data, run the models, investigate the results, and evaluate the next decision.
What I built
- The data foundation. Brought administration extracts, asset holdings, liability cash flows, market inputs, and model outputs into a common analytical environment. Built validation, standardization, and reconciliation workflows so users could inspect the inputs behind an analysis.
- The applications. Built and integrated the interfaces connecting data preparation, model execution, ALM metrics, attribution, and scenario exploration. Connected specialist calculation engines to applications that actuarial, investment, and finance teams could use together.
- The workflows. Connected inputs, calculation runs, dashboards, and saved results. Designed the experience around how users start an analysis, review a proposed action, investigate an exception, and return to a previous result.
- The AI layer. Built an analyst with access to the platform’s data, calculation logic, and execution tools. Users can move from a business question to supporting analysis and a new model run within the same conversation.
What the AI makes possible
- Explain what changed. Investigate movements in asset and liability values across periods, currencies, and business entities, with supporting tables, charts, and source references.
- Understand how a result was calculated. Trace a figure back through its data and methodology. Specialist agents investigate run data and calculation code, then combine their findings into a coherent explanation.
- Explore what happens next. Translate an interest-rate, currency, or assumption question into an analytical plan. Prepare changed inputs, launch an approved calculation, and compare the scenario with the baseline.
- Serve different audiences. Switch between executive explanations and technical analysis while retaining the context of the investigation.
The agent harness
I built the runtime around the agent: the services that give it context and tools, maintain the conversation, execute its work, and make that work visible to the user.
- Context and orchestration. Module-specific skills describe the platform’s data and calculation methods. Specialist agents divide investigations between data analysis and methodology tracing.
- Persistent execution. Warm agent processes reduce repeated startup work; saved session identifiers support conversation resumption. Streaming events expose analysis steps, tool activity, and results as they arrive.
- Operational controls. Interrupt controls, inactivity watchdogs, and cost limits give long-running investigations clear stopping and recovery paths. Model and reasoning-effort selection let users match the analysis to the task.
Evaluations and reliability
- Replayable AI evaluations. Built a seed set of grounded questions and a replay harness against the running backend, creating a repeatable basis for assessing answers as the agent evolves.
- Calculation reconciliation. Compared model outputs and attribution against reference calculations. In the demonstrated liability-attribution case, all reference rows reconciled.
- Workflow verification. Tested the complete path from a question and proposed plan through approval, model execution, result retrieval, and explanation, alongside backend regression checks and frontend builds.
- Feedback loops. Captured user feedback and exposed the steps behind an answer, making failures easier to investigate and turn into improvements to context, tools, and evaluation cases.
Guardrails and isolation
- Grounded answers. Source references and explicit data conventions anchor analysis to the correct period, units, and calculation basis. Analyst-derived estimates are identified separately from model-run results.
- Reviewable actions. Calculation plans appear as approval cards with progress and stop controls. Read-only SQL access supports data investigation; changed scenario inputs and outputs are kept separate from the baseline.
- Contained rendering and execution. Generated charts render in sandboxed frames. Model calculations run in separate child processes so a failed calculation can be recorded without taking down the API service.
Production architecture
I implemented the platform as a managed, multi-user service, with controls spanning identity, data access, agent execution, model delivery, and operations.
- Application and identity. The React interface and FastAPI services run on managed Kubernetes behind enterprise single sign-on. Role-based access is enforced for portfolios, datasets, and model runs at the API and data layers.
- Data and run management. Versioned analytical packages sit in object storage, with application metadata in MongoDB. Run records capture input versions, assumptions, code versions, approvals, and outputs so results can be reproduced and compared.
- Isolated agent execution. Analytical jobs run in short-lived containers with read-only baselines, separate writable workspaces, CPU and memory limits, and restricted network access. Scoped service credentials and explicit authorization control consequential actions.
- Model access and orchestration. An enterprise gateway manages model access, credentials, usage policies, and request telemetry. Long-running calculations are queued separately from interactive conversations, with cancellation and recoverable job states.
- Release evaluation. Versioned questions, reference calculations, and workflow tests gate releases. Checks cover numerical correctness, source grounding, tool selection, authorization boundaries, and recovery from missing data or failed runs.
- Operations and improvement. Correlated logs and traces track service health, run failures, latency, and model cost. Redacted feedback and reviewed failures feed the evaluation set, supported by staged releases and rollback.
Technology by layer
- Applications: React, TypeScript, Vite, AG Grid, Recharts.
- Services and calculations: Python, FastAPI, pandas, NumPy, SciPy, CVXPY.
- Analytical data: Parquet and DuckDB, with SQL access to inputs, model outputs, attribution, and saved scenarios.
- AI runtime: Claude Agent SDK and headless agent runtimes, specialist agents, domain skills, streamed tool events, persistent sessions, and an evaluation replay harness.
- Production infrastructure: Docker, managed Kubernetes, Okta/OIDC, MongoDB, object storage, a job queue, an enterprise model gateway, and OpenTelemetry.