I Built the Same Agent Twice to Find Out If the Harness Matters
A controlled comparison of a hand-rolled Python harness and Google ADK - and what happened when I audited my own headline
professional
Technical reports. Each one accompanies a repository - the write-up explains the work, the code is the proof.
A controlled comparison of a hand-rolled Python harness and Google ADK - and what happened when I audited my own headline
Does your AI agent actually know what it doesn't know? 3,800 controlled experiments across three frontier models, measuring whether the metadata around your data changes how reliable the agent is.
How classical time series, deep learning, and Generative AI come together to tackle one of supply chain's oldest problems — built on the M5 competition data, where 68% of item-level observations are zero.
Agentic Pharmacovigilance - automating adverse-event case processing and safety signal detection
Can You Trust an LLM to Grade Pharmaceutical Compliance? I Built a Framework to Find Out.