Skip to current window
Hoang Ha
Apps
ProjectsResearchWritingConferencesTravelAboutContactMeddies ↗
Writing

Notes from the work.

166 Gigabytes of Speech for a Language ASR Forgot cover2026-09-16, 7 min read166 Gigabytes of Speech for a Language ASR ForgotMeddies ASR Synthetic Dialog: doctor–patient consultations in Vietnamese, English, and Chinese — 22 scenario profiles, emotion that is spoken but never transcribed, and per-turn alignment for 15–30 second training windows.Teaching a Model to Consult, Not Just Answer cover2026-09-16, 7 min readTeaching a Model to Consult, Not Just AnswerMeddies Consultant: 109K English and 58K Vietnamese multi-turn clinical consultations built around how clinicians actually interview patients — Calgary-Cambridge, FIFE, OPQRST.Retrieval That Crosses Languages on Purpose cover2026-09-16, 7 min readRetrieval That Crosses Languages on PurposeThe design behind Meddies Embedding Data: query language drawn independently of passage language, every pair grounded and validated, every row carrying its provenance.22,336 Ways a Vietnamese Patient Can Break Your Model cover2026-09-16, 8 min read22,336 Ways a Vietnamese Patient Can Break Your ModelMeddies Patient Safety: a clinical red-team set probing five unsafe response modes with Vietnamese patient personas, folk-medicine phrasing, and family-driven questions English benchmarks never exercise.150,000 Vietnamese Patients Who Don't Exist cover2026-09-16, 8 min read150,000 Vietnamese Patients Who Don't ExistMeddies Persona: synthetic patient personas carrying province, dialect, economic tier, traditional-medicine use, and chief complaints in colloquial Vietnamese — because the population is the problem, not the add-on.Nine Labels, Seventeen Languages, One Human Review cover2026-09-16, 8 min readNine Labels, Seventeen Languages, One Human ReviewHow we built Meddies PII v2, a 350M span extractor for clinical de-identification, and what its benchmark numbers do and do not promise.2.9 Million Vietnamese Medical Questions, Grounded in the Literature cover2026-09-16, 6 min read2.9 Million Vietnamese Medical Questions, Grounded in the LiteratureMeddies QA: 2.94M chat-template QA rows and 7.11M question-only prompts across five medical domains — built from grounded sources, shipped with its limits attached.Seven Releases Before One Product cover2026-09-16, 5 min readSeven Releases Before One ProductA map of the Meddies research line: privacy, retrieval, QA, consultations, personas, safety red-teaming, and speech — what each artifact answers and what none of them claim.MeddiesAI: What In Progress Actually Means cover2026-09-16, 7 min readMeddiesAI: What In Progress Actually MeansBuilding clinical intelligence for Vietnamese hospitals — what exists, what doesn't yet, and why the research line came first.MedMeta: Testing LLMs on Evidence Synthesis, Not Recall cover2026-09-16, 4 min readMedMeta: Testing LLMs on Evidence Synthesis, Not RecallA benchmark built from 81 medical meta-analyses shows RAG workflows outperforming parametric knowledge, and a shared weakness: negated evidence slipping past every model tested.Pensez: French Reasoning with 2,000 Curated Examples cover2026-09-16, 3 min readPensez: French Reasoning with 2,000 Curated ExamplesA bilingual Qwen2.5 fine-tuning experiment, and why a small reasoning dataset is not the same thing as training a model on little data.Correcting Scientific Facts in Language Models: A Thesis Plan, Not a Thesis cover2026-09-16, 8 min readCorrecting Scientific Facts in Language Models: A Thesis Plan, Not a ThesisThree months into a PhD at Université Grenoble Alpes on explainable correction of scientific facts in LLMs — the question, why retractions, and what I do not know yet.Pretraining a Molecular Encoder on 98,973,909 SMILES cover2026-09-16, 3 min readPretraining a Molecular Encoder on 98,973,909 SMILESSELFIES for a RoBERTa-style molecular encoder: why Robust Molecular Representation, what a nearly 99-million-molecule pretraining set buys you, and the discipline of publishing before the results section exists.ToolMaestro: The Model That Ships With No Card cover2026-09-16, 3 min readToolMaestro: The Model That Ships With No CardA 7B-class tool-use model published as bare weights on Hugging Face — no model card, no benchmarks, no decoding of its name. What an empty card admits, and the problem the model targets: knowing when to call a tool at all.Vista: Building Vietnamese Image Descriptions as a Team cover2026-09-16, 3 min readVista: Building Vietnamese Image Descriptions as a TeamThe Vista dataset and Vistral-V-7B pair Vietnamese visual instruction data with SigLIP, a projector, and Vistral. Image description is the target, not OCR.Two Weekends, 23 Bugs, and a 144M Diffusion Language Model cover2026-03-11, 12 min readTwo Weekends, 23 Bugs, and a 144M Diffusion Language ModelFrom reading papers to training a 144M-parameter diffusion LLM on H100s. A deep dive into the bugs, optimizations, and lessons from building SmolDLM.How I Built a Personal Blog in 30 Minutes Using AI (No Code Needed) cover2025-06-14, 6 min readHow I Built a Personal Blog in 30 Minutes Using AI (No Code Needed)Want to create your own blog without writing code? Here's how I used an AI assistant and a tool called Lovable to get my personal website live - in under 30 minutes!How I Built a Personal Blog with LLMs (Part 2) cover2025-01-25, 8 min readHow I Built a Personal Blog with LLMs (Part 2)Join me on my journey from pharmacist to web developer, using AI as my buddy.How I Built a Personal Blog with LLMs (Part 1) cover2025-01-23, 7 min readHow I Built a Personal Blog with LLMs (Part 1)Join me on my journey from pharmacist to web developer, using AI as my buddy.