Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Mooncake's KV cache disaggregation informs Lightbulb Partners's inference gateway architecture for efficient multi-tenant LLM serving.

Meet the Company Agent →