LessUp/hpc-ai-optimization-lab

CUDA HPC Kernel Optimization Textbook: Naive to Tensor Core — GEMM, FlashAttention & Quantization | CUDA 高性能算子开发教科书：从 Naive 到 Tensor Core 完整优化路径，涵盖 GEMM/FlashAttention/量化

/ 100

Experimental

No Package No Dependents

Maintenance 13 / 25

Adoption 0 / 25

Maturity 9 / 25

Community 0 / 25

How are scores calculated?

Stars

—

Forks

—

Language

Cuda

License

MIT

Category

gpu-parallel-programming

Last pushed

Mar 13, 2026

Commits (30d)

GitHub

GPU Parallel Programming · 60 frameworks

Get this data via API

curl "https://pt-edge.onrender.com/api/v1/quality/ml-frameworks/LessUp/hpc-ai-optimization-lab"

Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.

Higher-rated alternatives

brucefan1983/GPUMD

Graphics Processing Units Molecular Dynamics

iree-org/iree

A retargetable MLIR-based machine learning compiler and runtime toolkit.

uxlfoundation/oneDAL

oneAPI Data Analytics Library (oneDAL)

rapidsai/cuml

cuML - RAPIDS Machine Learning Library

NVIDIA/cutlass

CUDA Templates and Python DSLs for High-Performance Linear Algebra

Explore ML Frameworks

All categories Trending ML Framework directory Insights