Tool Concierge
A LoRA fine-tune of Qwen3-8B that generates a single MCP tool call — or abstains — instead of hallucinating tool names and malformed arguments.
75% call accuracyup from 56% pretrain; abstention accuracy more than doubled
I build machine learning systems and try to understand why they work.
My work spans healthcare AI, audio/NLP, and lately mostly LLMs and agents. I'm less interested in whether a model produces an output than in why it behaves the way it does, where it fails, and whether the result deserves to be trusted.
A LoRA fine-tune of Qwen3-8B that generates a single MCP tool call — or abstains — instead of hallucinating tool names and malformed arguments.
75% call accuracyup from 56% pretrain; abstention accuracy more than doubled
Point it at YouTube videos and it transcribes, segments by topic, summarizes two ways, extracts entities, and graphs what recurs across videos — no paid APIs.
Zero paid APIsfaster-whisper + local Ollama, entities graphed in Neo4j
An LLM that turns raw emergency-department data into structured clinical risk summaries — the synthesis a physician does in their head between rooms.
118,385 encountersStanford MC-MED, scored for accuracy, hallucination, and equity