I automate infrastructure and build AI tools: LLM inference on a single RTX 5090, agents with Claude Code, and the Proxmox and Docker stack they run on.
Stack
AI
LLM inference, CUDA, C++, AI agents, Claude Code, Cline, ChatGPT
Automation
Ansible, CI/CD, Docker, Git, Bash, PowerShell
Virtualization
Proxmox, VMware vSphere, HA clustering, Dell PowerEdge, NetApp, Backup
Microsoft
Entra ID, Intune, Active Directory, Exchange, SCCM, SQL Server, SharePoint
LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.
Rejection is the default: an operating manual for a Claude Code session whose job is driving other sessions and refusing anything not backed by evidence
Local-first AI prompt generator with a Gradio UI: 132 tuned roles for video, image, audio, 3D and creative AI tools. Runs uncensored Dolphin LLMs fully offline, auto-scaled to your GPU VRAM. No API keys, no cloud.
Recover the CUDA launch and implicit-destructor edges a syntactic indexer loses, written straight into CodeGraph's SQLite graph. On imp (kekzl/imp, 240k lines of C++/CUDA): kernel launch coverage 74.2% → 99.3%, destructors 0% → 85%. On llm.c, which hides nothing, it adds 0 edges.
Monitor Entra ID (Azure AD) application secrets and certificates for expiration. Sends alerts via Email, Teams, Slack, or Webhook before credentials expire.
A living, continually learning neuromorphic being — a spiking neural network in C++23/CUDA with purely local plasticity (STDP, R-STDP, and Feedback Alignment instead of backprop). Runs on RTX 5090.