Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
gpu
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
High-Throughput LLM Inference & Training: A Deep Dive into vLLM
Aleksei Romanov
Aleksei Romanov
Aleksei Romanov
Follow
for
g factor
Sep 21
High-Throughput LLM Inference & Training: A Deep Dive into vLLM
#
ai
#
machinelearning
#
python
#
gpu
Comments
Add Comment
9 min read
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
Fyodor Serzhenko
Fyodor Serzhenko
Fyodor Serzhenko
Follow
Sep 18
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
#
cuda
#
gpu
#
performance
#
cpp
Comments
Add Comment
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
ComradePenguin
ComradePenguin
ComradePenguin
Follow
Sep 18
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
#
cuda
#
nvidia
#
gpu
1
 reaction
Comments
Add Comment
10 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Follow
Sep 17
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
#
machinelearning
#
pytorch
#
deepspeed
#
gpu
Comments
Add Comment
10 min read
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
Yehor Cherednichenko
Yehor Cherednichenko
Yehor Cherednichenko
Follow
Sep 17
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
#
machinelearning
#
performance
#
compilers
#
gpu
2
 reactions
Comments
1
 comment
6 min read
What Happens When You Ask an LLM a Question
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
What Happens When You Ask an LLM a Question
#
ai
#
localllm
#
llmbasics
#
gpu
Comments
Add Comment
8 min read
CUDA Cores vs Tensor Cores Explained
Sumukh Shenoy
Sumukh Shenoy
Sumukh Shenoy
Follow
Sep 14
CUDA Cores vs Tensor Cores Explained
#
gpu
#
machinelearning
#
deeplearning
#
beginners
Comments
1
 comment
2 min read
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
stmanst
stmanst
stmanst
Follow
Sep 9
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
#
ai
#
machinelearning
#
gpu
#
debugging
Comments
Add Comment
3 min read
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
erniou86
erniou86
erniou86
Follow
Sep 9
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
#
ai
#
gpu
#
selfhosting
#
machinelearning
Comments
1
 comment
3 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary
Stjepan
Stjepan
Stjepan
Follow
Sep 6
Reverse-Engineering NVIDIA: Modifying a CUDA binary
#
nvidia
#
gpu
#
cuda
#
hex
Comments
Add Comment
4 min read
From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Follow
Sep 5
From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn
#
ai
#
llm
#
gpu
#
machinelearning
Comments
Add Comment
17 min read
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Follow
Sep 5
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
#
ai
#
llm
#
gpu
#
machinelearning
Comments
Add Comment
13 min read
Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes
Ramkumar Nagaraj
Ramkumar Nagaraj
Ramkumar Nagaraj
Follow
Sep 4
Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes
#
kubernetes
#
autoscaling
#
gpu
#
devops
Comments
Add Comment
6 min read
DGX Spark (GB10) memory sizing for LLM serving: the numbers
Jahn
Jahn
Jahn
Follow
Sep 2
DGX Spark (GB10) memory sizing for LLM serving: the numbers
#
nvidia
#
llm
#
inference
#
gpu
Comments
Add Comment
7 min read
How Many AI Avatars Can One GPU Handle? Real-World Test Reveals 4 Avatars at ÂĄ7,600 Each per Month
orca_forge
orca_forge
orca_forge
Follow
Aug 30
How Many AI Avatars Can One GPU Handle? Real-World Test Reveals 4 Avatars at ÂĄ7,600 Each per Month
#
gpu
#
sre
#
docker
Comments
Add Comment
9 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account