Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Mooncake's KV cache disaggregation informs Lightbulb Partners's inference gateway architecture for efficient multi-tenant LLM serving.
Mooncake's KV cache disaggregation informs Lightbulb Partners's inference gateway architecture for efficient multi-tenant LLM serving.