01AIn progress
AI Evaluation · Engineering + Statistical Reasoning
AI-native research workflow
A growing suite of reusable skills and MCP tools for coordinating evidence, document production, review, and evaluation across a high-stakes statistical research workflow.
- The question
- How can agentic workflows support complex academic work while keeping sources, decisions, and human judgment inspectable?
- The artifact
- Completed components include reusable skills for styling statistical papers and evaluating presentation rehearsals; additional workflow and evaluation artifacts are in development.
Codex skillsMCPAgent workflowsEvaluation design
01BOpen source
Engineering · Statistical Reasoning + AI Evaluation
Modern Personalized Recommendation
A local, synthetic-data recommender pipeline that separates multi-channel retrieval, FM + DCN v2 ranking, a Transformer sequence tower, business reranking, and offline evaluation.
- The question
- What does a recommender look like when recall, ranking, policy, and cohort evaluation are treated as one system?
- The artifact
- Inspect the implemented local pipeline and its explicitly documented roadmap.
PyTorchHydraParquetDCN v2Transformers
01CHackathon
Engineering · AI Evaluation
Multi-agent pipeline
A hackathon prototype that removes text from product images, interprets visual content, generates Canadian English and French listings, and runs copy and language quality checks through a Streamlit workflow.
- The question
- How can a product-image workflow combine visual understanding, localization, generation, and quality control?
- The artifact
- Run the public Streamlit demo in mock mode or connect the documented inference workflow.
PythonStreamlitMultimodal modelsLLMs
01DHackathon
AI Evaluation · Engineering
Machine-generated text detector
A collaborative Microsoft Fabric hackathon project for classifying human and machine-generated language, with evaluation notebooks and documented robustness testing against adversarial transformations.
- The question
- Where does an efficient text detector generalize, and what happens when the input distribution is deliberately perturbed?
- The artifact
- Review the public model workflow, evaluation notebook, collaborators, and benchmark notes.
PythonTransformersLoRAE5-smallRAID