
1. 总览基于 Kubernetes 部署 FastAPI LLM Serving通过 Kubernetes Service 和 Ingress 对外提供 API并通过host.docker.internal访问 MacBook 上的 Ollama。核心调用链Client → Kubernetes Ingress → Kubernetes Service → FastAPI → host.docker.internal → MacBook → Ollama → Qwen3:8b → Response2. 构建 Docker Image2.1 第一次构建(.venv)(base)igwanglyangMacbookProdahouzicn GitHub %cdllm-serving(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving %dockerbuild-tllm-serving:0.1.0.构建完成[] Building 88.2s (11/11) FINISHED naming to docker.io/library/llm-serving:0.1.0检查镜像(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving %dockerimages|grepllm-servingllm-serving 0.1.0 51e808a2a4bc 10 seconds ago 200MB2.2 修改 Dockerfile 后重新构建修改 Dockerfile 后使用--no-cache(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving %dockerbuild --no-cache-tllm-serving:0.1.1.构建完成[] Building 38.9s (10/10) FINISHED naming to docker.io/library/llm-serving:0.1.13. Kubernetes Service 测试3.1 使用测试 Pod(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl run curl-test\-nllm-serving\--rm-it\--imagecurlimages/curl\--sh~ $curlhttp://llm-serving:8000/health{status:ok,service:llm-serving,ollama:ok}3.2 重新部署 Pod(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl delete pod-nllm-serving-lappllm-servingpod llm-serving-5d68446ff5-n9xn6 deleted(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get pods-nllm-serving-w最终llm-serving-5d68446ff5-xgs5l 1/1 Running4. Helm 部署(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % helminstallllm-serving ./k8sNAME: llm-serving NAMESPACE: default STATUS: deployed REVISION: 1检查资源(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get all-nllm-serving核心资源pod/llm-serving-5d68446ff5-n9xn6 1/1 Running service/llm-serving ClusterIP 10.106.230.103 8000/TCP deployment.apps/llm-serving 1/1 replicaset.apps/llm-serving 1/15. 检查 Kubernetes 资源(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get pods-nllm-serving(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get deployment-nllm-serving(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get configmap-nllm-serving核心状态Pod → 1/1 Running Deployment → 1/1 Available ConfigMap → llm-serving-config6. ConfigMap(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl describe configmap llm-serving-config-nllm-serving关键配置DEFAULT_MODEL: qwen3:8b OLLAMA_BASE_URL: http://host.docker.internal:11434调用关系FastAPI → host.docker.internal → MacBook → Ollama7. 测试 Pod 内 FastAPI(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectlexec-nllm-serving deploy/llm-serving --\curlhttp://localhost:8000/health当前镜像没有curlexec: curl: executable file not found in $PATH改用 Python(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectlexec-nllm-serving deploy/llm-serving --\python-cimport urllib.request; print(urllib.request.urlopen(http://localhost:8000/health).read().decode()){status:ok,service:llm-serving,ollama:ok}8. 测试 Pod → Ollama(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectlexec-nllm-serving deploy/llm-serving --\python-cimport urllib.request; print(urllib.request.urlopen(http://host.docker.internal:11434/api/tags).read().decode())确认 Ollama 返回模型列表其中包含qwen3:8b核心网络Kubernetes Pod → host.docker.internal → MacBook → Ollama9. 验证 FastAPI 配置9.1 Health(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectlexec-nllm-serving deploy/llm-serving --\python-cimport urllib.request; print(urllib.request.urlopen(http://localhost:8000/health).read().decode()){status:ok,service:llm-serving,ollama:ok}9.2 Environment(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectlexec-nllm-serving deploy/llm-serving --\env|grep-EOLLAMA|DEFAULT_MODELOLLAMA_BASE_URLhttp://host.docker.internal:11434 DEFAULT_MODELqwen3:8b10. Kubernetes Service(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get svc-nllm-servingNAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) llm-serving ClusterIP 10.106.230.103 none 8000/TCP11. MacBook → Kubernetes Service11.1 Port Forward(.venv)(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl port-forward\-nllm-serving\service/llm-serving\8000:8000Forwarding from 127.0.0.1:8000 - 8000 Forwarding from [::1]:8000 - 8000 Handling connection for 800011.2 测试另一个窗口(base)igwanglyangMacbookProdahouzicn ~ %curlhttp://localhost:8000/health{status:ok,service:llm-serving,ollama:ok}12. Kubernetes → Ollama → Qwen3(base)igwanglyangMacbookProdahouzicn ~ %curlhttp://localhost:8000/v1/chat/completions\-HContent-Type: application/json\-d{ model: qwen3:8b, messages: [ { role: user, content: Explain Kubernetes Service in 100 words. } ] }返回{id:chatcmpl-6a1f8e7d9efd,object:chat.completion,created:1791474852,model:qwen3:8b,choices:[{index:0,message:{role:assistant,content:A Kubernetes Service is an abstraction layer that provides stable network access to a group of pods (or a single pod) through a consistent IP address and DNS name. It defines policies for load balancing, traffic routing, and discovery, enabling reliable communication within a cluster or externally. Services act as logical endpoints, abstracting the dynamic IP changes of pods. They support types like ClusterIP (internal), NodePort (exposes on node ports), and LoadBalancer (external). By decoupling networking from pod lifecycle, Services ensure scalability, resilience, and simplified service discovery, making microservices and distributed systems more manageable in Kubernetes environments. (100 words)},finish_reason:stop}],usage:{prompt_tokens:21,completion_tokens:371,total_tokens:392}}13. Pod 自愈(base)igwanglyangMacbookProdahouzicn ~ % kubectl delete pod\-nllm-serving\-lappllm-servingpod llm-serving-5d68446ff5-xgs5l deleted观察(base)igwanglyangMacbookProdahouzicn llm-serving % kubectl get pods-nllm-serving-w最终llm-serving-5d68446ff5-ppd5d 1/1 Running14. Deployment 与 Health Probe14.1 Deployment(base)igwanglyangMacbookProdahouzicn ~ % kubectl get deployment-nllm-servingllm-serving 1/1 1 114.2 Pod(base)igwanglyangMacbookProdahouzicn ~ % kubectl describe pod\-nllm-serving\-lappllm-serving关键状态Status: Running Node: docker-desktop/192.168.65.3 IP: 10.1.19.16614.3 Health ProbeLivenessLiveness → 是否需要重启ReadinessReadiness → 是否可以接收流量当前配置Liveness: http-get http://:8000/health delay15s timeout1s period20s #success1 #failure3 Readiness: http-get http://:8000/health delay5s timeout1s period10s #success1 #failure315. Ingress15.1 检查 Ingress(base)igwanglyangMacbookProdahouzicn ~ % kubectl get svc-A|grep-Eingress|harbor关键结果ingress-nginx-controller NodePort 10.104.86.242 none 80:30080/TCP,443:30443/TCP检查 Pod(base)igwanglyangMacbookProdahouzicn ~ % kubectl get pods-A|grep-Eingress|harboringress-nginx-controller-df569d69f-q59z2 1/1 Running15.2 测试 Ingress(base)igwanglyangMacbookProdahouzicn ~ %curl-HHost: llm-serving.localhttp://127.0.0.1:30080/health{status:ok,service:llm-serving,ollama:ok}调用链MacBook → NodePort 30080 → Ingress → Service → FastAPI16. LLM API(base)igwanglyangMacbookProdahouzicn ~ %curlhttp://llm-serving.local:30080/v1/chat/completions\-HContent-Type: application/json\-d{ model: qwen3:8b, messages: [ { role: user, content: Explain Kubernetes Ingress in 50 words. } ] }返回{id:chatcmpl-f90f3c2f9019,object:chat.completion,created:1791477303,model:qwen3:8b,choices:[{index:0,message:{role:assistant,content:Kubernetes Ingress is a resource that manages external access to services in a cluster, routing HTTP/HTTPS traffic based on rules. It enables load balancing, SSL termination, and exposes services via a single IP, using an Ingress controller to handle the actual routing logic.},finish_reason:stop}],usage:{prompt_tokens:21,completion_tokens:205,total_tokens:226}}17. Rolling UpdateRolling Update Kubernetes 的无停机/低中断版本发布机制。Old Pod → New Pod → Traffic gradually moves → Old Pod removed验证kubectl get deployment-nllm-serving kubectl get pods-nllm-serving-wcurlhttp://llm-serving.local:30080/health18. HelmHelm 可以理解成Kubernetes 的 Package Manager Deployment Template Engine。当前部署helminstallllm-serving ./k8s调用关系Helm Chart → Kubernetes Manifest → Deployment → ReplicaSet → Pod → Service → Ingress最终架构Client → Kubernetes Ingress → Kubernetes Service → FastAPI → host.docker.internal → MacBook → Ollama → Qwen3:8b → Response完成 Kubernetes → FastAPI → Ollama → Qwen3 的完整 LLM Serving 链路验证。