Contributor to the open-source community llama.cpp
Pinned Loading
-
Inference-Optimization
Inference-Optimization PublicInference-Optimization — CUDA inference optimization experiment based on C++ LLM Runtime For the Qwen2/LLaMA model, complete performance analysis, operator optimization, and regression verification…
C++ 1
-
MCP_Server
MCP_Server PublicHigh-performance C++ MCP server using stdio transport and JSON-RPC 2.0 protocol. Supports Content-Length framing, exception handling, and graceful exit for IPC, AI tools, and plugin systems.
C++
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.
