Efficient LLM inference on Slurm clusters.
-
Updated
Jul 22, 2026 - Python
Efficient LLM inference on Slurm clusters.
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。
Multi-model AI agent runtime. Define agents in YAML, connect 6 LLM providers, orchestrate with ReAct/Plan&Execute/Fan-Out/Pipeline/Supervisor/Swarm patterns, and deploy as REST/WebSocket API with RAG, memory, MCP tools, guardrails, and OpenTelemetry observability.
AI Workload Control Layer for routing deterministic, reusable, retrieval-needed, tool-needed, and provider-needed work before model invocation.
Reliability control for the Anthropic Python SDK
A browser-based UI for launching, monitoring, clustering and managing multiple llama.cpp server instances from inside a Docker container. Includes an Ollama-compatible API proxy
Persistent Cognitive Memory Infrastructure — durable, multi-tenant memory layer for AI agents. HTTP + gRPC, hybrid retrieval, background workers, observability. Go.
A production-grade, schema-aware PostgreSQL MCP server for enterprise AI. Features Zero-Trust SQL validation, multi-tier permissions, and real-time schema introspection for secure, autonomous database operations.
Krako 2.0 – Energy-efficient, triadic multi-tier inference infrastructure enabling adaptive routing across heterogeneous edge–cloud nodes.
Joule is a budget-controlled AI agent runtime that minimizes energy and token usage through hierarchical routing and deterministic tool execution.
An intelligent gateway for Claude APIs that dynamically routes requests to the most cost-efficient model, caches responses, and escalates based on confidence signals — reducing LLM spend without sacrificing quality.
secrets.wtf: defensive AI infrastructure exposure index for Ollama and LM Studio hosts, local LLM APIs, model observations, remediation, and takedown requests.
A Branchless, Zero-Jitter Ingress Router for 32-GPU Distributed Mesh Networks utilizing JAX/XLA and NCCL.
Enterprise-grade Sovereign AI Stack optimized for NVIDIA Blackwell (sm_120) & vLLM. Features 256K context window, 5.8k tok/s prefill, and integrated observability via Langfuse.
Deterministic state engine for managing conversation state and constraints in LLM applications.
One command. Full LLM stack. Zero config.
Multi-Agent AI Orchestrator 2026 🚀 | YAML, 6+ LLM Providers, ReAct & Swarm
A lightweight Bun + Express template that connects to the Testune AI API and streams chat responses in real time using Server-Sent Events (SSE)
Add a description, image, and links to the llm-infrastructure topic page so that developers can more easily learn about it.
To associate your repository with the llm-infrastructure topic, visit your repo's landing page and select "manage topics."