Groq
Ultra-fast inference API for open models on custom LPU hardware
PyTorch Foundation
High-throughput serving engine for running open LLMs with an OpenAI-compatible API