Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Type
/
to search
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
handshape
/
llama-cpp-python
Public
forked from
abetlen/llama-cpp-python
Notifications
You must be signed in to change notification settings
Fork
0
Star
0
Code
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Pull requests
Actions
Projects
Security and quality
Insights
Commits
Branch selector
patch-2
User selector
All users
All time
Commit history
Commits on Jun 1, 2026
docs: reference transformers chat template extensions
abetlen
committed
3abd8fa
View commit details
Copy full SHA for 3abd8fa
Browse repository at this point
Merge branch 'main' into patch-2
abetlen
committed
2c5f893
View commit details
Copy full SHA for 2c5f893
Browse repository at this point
Commits on May 31, 2026
fix: avoid cleanup errors for partially initialized LlamaModel (#2173)
Show description for fdf38b3
usernames122
and
abetlen
authored
fdf38b3
View commit details
Copy full SHA for fdf38b3
Browse repository at this point
fix: suppress stdout and stderr in Jupyter notebooks (#2181)
Show description for 6bdab5d
Anai-Guo
authored
6bdab5d
View commit details
Copy full SHA for 6bdab5d
Browse repository at this point
feat: enable arm64 musl builds (#2221)
Show description for b91460b
acon96
and
abetlen
authored
b91460b
View commit details
Copy full SHA for b91460b
Browse repository at this point
Fix: model fails to load when chat template uses HuggingFace generation tags (#2226)
Show description for f160bf7
tobocop2
and
abetlen
authored
f160bf7
View commit details
Copy full SHA for f160bf7
Browse repository at this point
feat: Update llama.cpp to d749821db (#2233)
abetlen
authored
2c455a5
View commit details
Copy full SHA for 2c455a5
Browse repository at this point
Commits on May 24, 2026
docs: add contributing guide (#2229)
abetlen
authored
3bda091
View commit details
Copy full SHA for 3bda091
Browse repository at this point
Commits on May 23, 2026
feat: Update llama.cpp to c0c7e147e (#2228)
abetlen
authored
52fe54b
View commit details
Copy full SHA for 52fe54b
Browse repository at this point
Commits on May 18, 2026
feat: Update llama.cpp to b9a2170fc (#2223)
abetlen
authored
5dd9b1c
View commit details
Copy full SHA for 5dd9b1c
Browse repository at this point
Commits on May 15, 2026
chore: migrate llama.cpp submodule to ggml-org (#2034)
Show description for c7bea71
shalinib-ibm
and
abetlen
authored
c7bea71
View commit details
Copy full SHA for c7bea71
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@91e84fed6 (#2218)
Show description for 7664a3e
abetlen
authored
7664a3e
View commit details
Copy full SHA for 7664a3e
Browse repository at this point
Commits on May 13, 2026
fix(embedding): set kv_unified=True when embedding=True to enable batch processing (#2217)
Show description for 95ccb19
SanjanaB123
and
abetlen
authored
95ccb19
View commit details
Copy full SHA for 95ccb19
Browse repository at this point
Commits on May 11, 2026
chore: bump version to 0.3.23 (#2215)
abetlen
authored
4a1a8ec
View commit details
Copy full SHA for 4a1a8ec
Browse repository at this point
feat: update llama.cpp to 7d442abf (#2214)
abetlen
authored
5684112
View commit details
Copy full SHA for 5684112
Browse repository at this point
fix(embed): mark all tokens as output to suppress llama.cpp 'overriding' INFO (#2208) (#2212)
Anai-Guo
authored
f8c1f36
View commit details
Copy full SHA for f8c1f36
Browse repository at this point
Commits on May 8, 2026
feat: update llama.cpp to 5d6f18a63 (#2207)
abetlen
authored
f774690
View commit details
Copy full SHA for f774690
Browse repository at this point
fix: configure n_seq_max for batched embeddings (#2206)
Show description for 128c331
abetlen
authored
128c331
View commit details
Copy full SHA for 128c331
Browse repository at this point
Commits on May 4, 2026
fix(_internals): use n_tokens0 offset when enabling last-token logits in add_sequence (#2205)
Show description for 90e8df9
Anai-Guo
authored
90e8df9
View commit details
Copy full SHA for 90e8df9
Browse repository at this point
Commits on May 2, 2026
fix(ci): skip unsupported Windows CUDA versions (#2204)
abetlen
authored
14d7846
View commit details
Copy full SHA for 14d7846
Browse repository at this point
fix(ci): install CUDA CCCL headers for wheel builds (#2203)
abetlen
authored
bc6ff9f
View commit details
Copy full SHA for bc6ff9f
Browse repository at this point
fix(ci): pass CUDA compiler arg for Windows detection (#2202)
abetlen
authored
04a3638
View commit details
Copy full SHA for 04a3638
Browse repository at this point
fix(ci): pass CUDA unsupported compiler flag during detection (#2201)
abetlen
authored
2bfd80c
View commit details
Copy full SHA for 2bfd80c
Browse repository at this point
chore: bump version to 0.3.22 (#2200)
abetlen
authored
9cf0ce7
View commit details
Copy full SHA for 9cf0ce7
Browse repository at this point
feat(ci): re-enable Windows CUDA wheels (#2198)
Show description for d2113a1
abetlen
authored
d2113a1
View commit details
Copy full SHA for d2113a1
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@63d93d173 (#2197)
abetlen
authored
587d94a
View commit details
Copy full SHA for 587d94a
Browse repository at this point
Commits on Apr 27, 2026
fix(docs): update mkdocstrings inventories config (#2195)
abetlen
authored
c6dc905
View commit details
Copy full SHA for c6dc905
Browse repository at this point
fix(ci): Scope CPU release wheel selectors by OS (#2194)
abetlen
authored
d2bcbac
View commit details
Copy full SHA for d2bcbac
Browse repository at this point
fix(ci): Repair py3 CPU release wheels (#2193)
abetlen
authored
195cc59
View commit details
Copy full SHA for 195cc59
Browse repository at this point
chore: bump version to 0.3.21 (#2192)
abetlen
authored
c8075d1
View commit details
Copy full SHA for c8075d1
Browse repository at this point
fix(ci): Build one arm64 py3 release wheel (#2191)
Show description for 511b3f4
abetlen
authored
511b3f4
View commit details
Copy full SHA for 511b3f4
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@f53577432 (#2189)
Show description for d87bf08
abetlen
authored
d87bf08
View commit details
Copy full SHA for d87bf08
Browse repository at this point
Commits on Apr 13, 2026
feat: Update llama.cpp to ggerganov/llama.cpp@227ed28e1 (#2182)
abetlen
authored
1b1a320
View commit details
Copy full SHA for 1b1a320
Browse repository at this point
Commits on Apr 8, 2026
feat: Update llama.cpp to ggerganov/llama.cpp@3bd9aa1f9 (#2176)
Show description for 1bcc5bc
abetlen
authored
1bcc5bc
View commit details
Copy full SHA for 1bcc5bc
Browse repository at this point
Commits on Apr 3, 2026
chore: bump version to 0.3.20 (#2171)
abetlen
authored
02d6bee
View commit details
Copy full SHA for 02d6bee
Browse repository at this point
Previous
Next
You can’t perform that action at this time.