Skip to content

kekz

I automate infrastructure and build AI tools: LLM inference on a single RTX 5090, agents with Claude Code, and the Proxmox and Docker stack they run on.

Stack

AI
LLM inference, CUDA, C++, AI agents, Claude Code, Cline, ChatGPT
Automation
Ansible, CI/CD, Docker, Git, Bash, PowerShell
Virtualization
Proxmox, VMware vSphere, HA clustering, Dell PowerEdge, NetApp, Backup
Microsoft
Entra ID, Intune, Active Directory, Exchange, SCCM, SQL Server, SharePoint
Linux services
nginx, Postfix, PostgreSQL, MySQL, NFS, Samba
Network and security
Firewalls, VLAN, WLAN, Fiber networks, Vulnerability analysis, Penetration testing

Projects

Public repositories on GitHub, most recently pushed first.

imp

LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.

Language
Cuda
Stars
43
Last push
Updated