Zanele Dlamini
zanele-d02
ยท
AI & ML interests
Efficient LLM inference, KV cache optimization, quantization, speculative decoding, model pruning
Recent Activity
liked a model about 10 hours ago
inference4j/efficientnet-lite4 upvoted a paper about 10 hours ago
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks upvoted a paper about 10 hours ago
VoxMem: Benchmarking Multimodal Memory in Large Audio Language ModelsOrganizations
None yet