English | ๐ŸŒ ็น้ซ”ไธญๆ–‡

Hi, I'm Sean Chuang (่ŽŠๅฎถ่ณ“), an engineer focused on building and shipping AI applications. Below are projects I built end-to-end โ€” model training, cloud deployment, and live demos. Feel free to click in and try them.

๐Ÿง‘โ€๐Ÿ’ป Tech Stack Overview

Category Technologies
AI / Self-hosted LLM ยท VLM Gemma-4-26B, Qwen3.6, Qwen3-VL, MedGemma, Llama-3.2-Vision, Qwen2.5-7B / 14B, llama.cpp / llama-server, vLLM,TensorRT-LLM, GGUF / 4-bit quantization, AWQ / GPTQ / FP8(ModelOpt)
Agent Architecture Planner / Executor (ReAct) / Verifier, MCP (Model Context Protocol)
RAG / Retrieval BM25 (tuned), bge-m3 (dense), RRF hybrid fusion, bge-reranker-v2-m3, NVIDIA NIM reranker
LLM Evaluation LLM-as-judge (faithfulness), Cohen's ฮบ calibration, stratified calibration sets, recall@k / hit@k / MRR, two-way human audit of the judge
Trustworthy AI / Uncertainty Conformal Prediction (statistical guarantees), Selective Prediction (abstain & defer to humans), cross-model disagreement, Riskโ€“Coverage / per-label AUC
Deep Learning Frameworks PyTorch, TensorFlow, ONNX, HuggingFace transformers
Computer Vision ViT, EfficientNetV2 + CBAM, YOLOv8 Nano (Ultralytics), OpenCV, TorchXRayVision
Software Quality / Testing pytest, hypothesis (property-based), mutation testing, mypy strict, ruff, pip-audit, GitHub Actions CI
Training & Explainability Focal Loss, Mixup, CutMix, Grad-CAM, Score-CAM, SHAP, Confusion Matrix
Healthcare Data Standards HAPI FHIR R4, MedAgentBench, MIMIC-CXR / CheXpert
Backend / API FastAPI, Flask, Nginx, Uvicorn, Gradio
Cloud / Deployment AWS EC2 / RDS, GCP Compute Engine / Cloud Run, Docker
System Integration LineBot, MySQL, HTML / JS
Tools / Languages MobaXterm, Git, Python, WSL2

๐Ÿ“‚ Project 2: Choosing a RAG Configuration on a Single GPU (SelectRAG)

๐Ÿ† Key Results

๐Ÿ‘‰ Answers "which RAG setup should we run?" with controlled experiments on a single RTX 4090: retrieval, inference engine and numeric format are changed one axis at a time. Hybrid + reranker lifts recall@3 from 0.808 to 0.941, and GPTQ-Int4 serves 2.36ร— the throughput of bf16. Even the LLM judge used for scoring was hand-audited in both directions, confirming quantization added no measurable errors.

๐Ÿ”— Full technical details (GitHub Pages)

๐Ÿ”— Source code (GitHub)

โ–ถ๏ธ Watch the Web Demo

๐Ÿงฐ Technical Highlights