Build the system.
Evaluate the behavior.

Recommendation, ranking, multimodal workflows, and AI evaluation treated as connected system-design problems—not isolated model demos.

01AIn progress

AI Evaluation · Engineering + Statistical Reasoning

AI-native research workflow

A growing suite of reusable skills and MCP tools for coordinating evidence, document production, review, and evaluation across a high-stakes statistical research workflow.

The question
How can agentic workflows support complex academic work while keeping sources, decisions, and human judgment inspectable?
The artifact
Completed components include reusable skills for styling statistical papers and evaluating presentation rehearsals; additional workflow and evaluation artifacts are in development.
stat-paper-styling-skillpresentation-rehearsal-feedback-skill
Codex skillsMCPAgent workflowsEvaluation design
01BOpen source

Engineering · Statistical Reasoning + AI Evaluation

Modern Personalized Recommendation

A local, synthetic-data recommender pipeline that separates multi-channel retrieval, FM + DCN v2 ranking, a Transformer sequence tower, business reranking, and offline evaluation.

The question
What does a recommender look like when recall, ranking, policy, and cohort evaluation are treated as one system?
The artifact
Inspect the implemented local pipeline and its explicitly documented roadmap.
Repository
PyTorchHydraParquetDCN v2Transformers
01CHackathon

Engineering · AI Evaluation

Multi-agent pipeline

A hackathon prototype that removes text from product images, interprets visual content, generates Canadian English and French listings, and runs copy and language quality checks through a Streamlit workflow.

The question
How can a product-image workflow combine visual understanding, localization, generation, and quality control?
The artifact
Run the public Streamlit demo in mock mode or connect the documented inference workflow.
Repository
PythonStreamlitMultimodal modelsLLMs
01DHackathon

AI Evaluation · Engineering

Machine-generated text detector

A collaborative Microsoft Fabric hackathon project for classifying human and machine-generated language, with evaluation notebooks and documented robustness testing against adversarial transformations.

The question
Where does an efficient text detector generalize, and what happens when the input distribution is deliberately perturbed?
The artifact
Review the public model workflow, evaluation notebook, collaborators, and benchmark notes.
RepositoryDemo ↗
PythonTransformersLoRAE5-smallRAID