English | ๐ ็น้ซไธญๆ
Hi, I'm Sean Chuang (่ๅฎถ่ณ), an engineer focused on building and shipping AI applications. Below are projects I built end-to-end โ model training, cloud deployment, and live demos. Feel free to click in and try them.
๐งโ๐ป Tech Stack Overview
| Category | Technologies |
|---|---|
| AI / Self-hosted LLM ยท VLM | Gemma-4-26B, Qwen3.6, Qwen3-VL, MedGemma, Llama-3.2-Vision, Qwen2.5-7B / 14B, llama.cpp / llama-server, vLLM,TensorRT-LLM, GGUF / 4-bit quantization, AWQ / GPTQ / FP8(ModelOpt) |
| Agent Architecture | Planner / Executor (ReAct) / Verifier, MCP (Model Context Protocol) |
| RAG / Retrieval | BM25 (tuned), bge-m3 (dense), RRF hybrid fusion, bge-reranker-v2-m3, NVIDIA NIM reranker |
| LLM Evaluation | LLM-as-judge (faithfulness), Cohen's ฮบ calibration, stratified calibration sets, recall@k / hit@k / MRR, two-way human audit of the judge |
| Trustworthy AI / Uncertainty | Conformal Prediction (statistical guarantees), Selective Prediction (abstain & defer to humans), cross-model disagreement, RiskโCoverage / per-label AUC |
| Deep Learning Frameworks | PyTorch, TensorFlow, ONNX, HuggingFace transformers |
| Computer Vision | ViT, EfficientNetV2 + CBAM, YOLOv8 Nano (Ultralytics), OpenCV, TorchXRayVision |
| Software Quality / Testing | pytest, hypothesis (property-based), mutation testing, mypy strict, ruff, pip-audit, GitHub Actions CI |
| Training & Explainability | Focal Loss, Mixup, CutMix, Grad-CAM, Score-CAM, SHAP, Confusion Matrix |
| Healthcare Data Standards | HAPI FHIR R4, MedAgentBench, MIMIC-CXR / CheXpert |
| Backend / API | FastAPI, Flask, Nginx, Uvicorn, Gradio |
| Cloud / Deployment | AWS EC2 / RDS, GCP Compute Engine / Cloud Run, Docker |
| System Integration | LineBot, MySQL, HTML / JS |
| Tools / Languages | MobaXterm, Git, Python, WSL2 |
๐ Project 2: Choosing a RAG Configuration on a Single GPU (SelectRAG)
๐ Key Results
๐ Answers "which RAG setup should we run?" with controlled experiments on a single RTX 4090: retrieval, inference engine and numeric format are changed one axis at a time. Hybrid + reranker lifts recall@3 from 0.808 to 0.941, and GPTQ-Int4 serves 2.36ร the throughput of bf16. Even the LLM judge used for scoring was hand-audited in both directions, confirming quantization added no measurable errors.
๐ Full technical details (GitHub Pages)
๐งฐ Technical Highlights