Hippocratic AI is seeking an LLM Inference Engineer to own and optimize the serving infrastructure for its healthcare AI platform, focusing on reducing latency, improving throughput, and cutting costs for millions of patient conversations. The role involves designing distributed serving architectures, implementing quantization and speculative decoding, and working with Python, C++, and CUDA.