Technical essays on agent infrastructure, evaluation of frontier models in vertical domains, retrieval-augmented generation, forecasting, and causal inference. Most were written for the Tickr engineering blog; where I could confirm the published URL, each post links back to the original.
An agent is only as useful as the context it can assemble, the tools it can invoke, and the constraints that govern how it acts. What it takes to turn a model into a system.
Generative Predictor Search: LLM reasoning to propose and validate covariates. 15.6% average MAPE reduction, 13× faster, with logical validation of predictors.
Why sentiment is a poor proxy for exposure, and what an event-linked risk index looks like instead.
Frontier models generate risk indices quickly, but reliable measurement needs transparent, event-linked signals. A head-to-head.
Dynamic hierarchical product categorization: a domain-tuned system beating OpenAI and Anthropic, at 100–1,000× human speed.
How to tune a RAG system when you have no labelled data yet: chunking, retrieval parameters, and the MMR diversity trade-off.
Year-over-year comparison is the wrong tool when the data-generating process isn't constant. Building a counterfactual instead. (Basis for a US patent application.)
Four things to settle before expanding investment in generative AI: data sourcing, security, model selection, and evaluation.