Inference Engineer · Scale AI
Powered by
CUDA / GPU Kernels
GitHub
Sparse Attention
arXiv + GitHub
KV Cache Optimization
GitHub + OpenReview
Distributed Training
GitHub
Quantization (FP8/INT4)
Papers
Inference Serving
GitHub
RLHF / Fine-tuning
OpenReview
Triton Kernels
GitHub
3 merged PRs — PagedAttention perf, chunked prefill memory
Efficient KV Cache Eviction via Graph-Structured Importance Scoring
Issue #812 — sliding window + GQA correctness fix
2 papers reviewed — sparse MoE routing, long-context eval
Speculative Decoding with Draft Token Pruning
Custom attention plugin for GQA + sliding window
Engineers most likely to collaborate with Alex based on graph proximity, shared contributions, and topic overlap.
Research Eng · DeepMind · 1.2mi
3 co-cited papers on KV eviction; both contributed to vLLM in the last 90 days.
PhD · Stanford NLP · 4.1mi
Sparse attention co-author; working on similar long-context efficiency problems for LLM serving.
MLSys · Together AI · 2.8mi
Active on flash-attn and vLLM; similar CUDA kernel optimization background.
Powered by graph traversal + embedding similarity