Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Type
/
to search
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
handshape
/
llama-cpp-python
Public
forked from
abetlen/llama-cpp-python
Notifications
You must be signed in to change notification settings
Fork
0
Star
0
Code
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Pull requests
Actions
Projects
Security and quality
Insights
Commits
Breadcrumbs
History for
llama-cpp-python
llama_cpp
on
patch-2
User selector
All users
All time
Commit history
Commits on Jun 1, 2026
docs: reference transformers chat template extensions
abetlen
committed
3abd8fa
View commit details
Copy full SHA for 3abd8fa
View code at this point
Browse repository at this point
Merge branch 'main' into patch-2
abetlen
committed
2c5f893
View commit details
Copy full SHA for 2c5f893
View code at this point
Browse repository at this point
Commits on May 31, 2026
fix: avoid cleanup errors for partially initialized LlamaModel (#2173)
Show description for fdf38b3
usernames122
and
abetlen
authored
fdf38b3
View commit details
Copy full SHA for fdf38b3
View code at this point
Browse repository at this point
fix: suppress stdout and stderr in Jupyter notebooks (#2181)
Show description for 6bdab5d
Anai-Guo
authored
6bdab5d
View commit details
Copy full SHA for 6bdab5d
View code at this point
Browse repository at this point
Fix: model fails to load when chat template uses HuggingFace generation tags (#2226)
Show description for f160bf7
tobocop2
and
abetlen
authored
f160bf7
View commit details
Copy full SHA for f160bf7
View code at this point
Browse repository at this point
Commits on May 18, 2026
feat: Update llama.cpp to b9a2170fc (#2223)
abetlen
authored
5dd9b1c
View commit details
Copy full SHA for 5dd9b1c
View code at this point
Browse repository at this point
Commits on May 15, 2026
feat: Update llama.cpp to ggerganov/llama.cpp@91e84fed6 (#2218)
Show description for 7664a3e
abetlen
authored
7664a3e
View commit details
Copy full SHA for 7664a3e
View code at this point
Browse repository at this point
Commits on May 13, 2026
fix(embedding): set kv_unified=True when embedding=True to enable batch processing (#2217)
Show description for 95ccb19
SanjanaB123
and
abetlen
authored
95ccb19
View commit details
Copy full SHA for 95ccb19
View code at this point
Browse repository at this point
Commits on May 11, 2026
chore: bump version to 0.3.23 (#2215)
abetlen
authored
4a1a8ec
View commit details
Copy full SHA for 4a1a8ec
View code at this point
Browse repository at this point
fix(embed): mark all tokens as output to suppress llama.cpp 'overriding' INFO (#2208) (#2212)
Anai-Guo
authored
f8c1f36
View commit details
Copy full SHA for f8c1f36
View code at this point
Browse repository at this point
Commits on May 8, 2026
feat: update llama.cpp to 5d6f18a63 (#2207)
abetlen
authored
f774690
View commit details
Copy full SHA for f774690
View code at this point
Browse repository at this point
fix: configure n_seq_max for batched embeddings (#2206)
Show description for 128c331
abetlen
authored
128c331
View commit details
Copy full SHA for 128c331
View code at this point
Browse repository at this point
Commits on May 4, 2026
fix(_internals): use n_tokens0 offset when enabling last-token logits in add_sequence (#2205)
Show description for 90e8df9
Anai-Guo
authored
90e8df9
View commit details
Copy full SHA for 90e8df9
View code at this point
Browse repository at this point
Commits on May 2, 2026
chore: bump version to 0.3.22 (#2200)
abetlen
authored
9cf0ce7
View commit details
Copy full SHA for 9cf0ce7
View code at this point
Browse repository at this point
Commits on Apr 27, 2026
chore: bump version to 0.3.21 (#2192)
abetlen
authored
c8075d1
View commit details
Copy full SHA for c8075d1
View code at this point
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@f53577432 (#2189)
Show description for d87bf08
abetlen
authored
d87bf08
View commit details
Copy full SHA for d87bf08
View code at this point
Browse repository at this point
Commits on Apr 13, 2026
feat: Update llama.cpp to ggerganov/llama.cpp@227ed28e1 (#2182)
abetlen
authored
1b1a320
View commit details
Copy full SHA for 1b1a320
View code at this point
Browse repository at this point
Commits on Apr 8, 2026
feat: Update llama.cpp to ggerganov/llama.cpp@3bd9aa1f9 (#2176)
Show description for 1bcc5bc
abetlen
authored
1bcc5bc
View commit details
Copy full SHA for 1bcc5bc
View code at this point
Browse repository at this point
Commits on Apr 3, 2026
chore: bump version to 0.3.20 (#2171)
abetlen
authored
02d6bee
View commit details
Copy full SHA for 02d6bee
View code at this point
Browse repository at this point
fix(misc): replace deprecated llama.cpp references (#2170)
Show description for 08e088c
abetlen
authored
08e088c
View commit details
Copy full SHA for 08e088c
View code at this point
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@f49e9178767d557a522618b16ce8694f9ddac628 (#2169)
abetlen
authored
100b275
View commit details
Copy full SHA for 100b275
View code at this point
Browse repository at this point
Commits on Mar 30, 2026
feat(server): add model-load chat_template_kwargs (#2168)
abetlen
authored
7257ba9
View commit details
Copy full SHA for 7257ba9
View code at this point
Browse repository at this point
Commits on Mar 25, 2026
Bump version to 0.3.19 (#2162)
abetlen
authored
f54421b
View commit details
Copy full SHA for f54421b
View code at this point
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@c0159f9c1f874da15e94f371d136f5920b4b5335 (#2161)
Show description for c670222
abetlen
authored
c670222
View commit details
Copy full SHA for c670222
View code at this point
Browse repository at this point
fix: handle embedding models without KV memory (#2160)
Show description for ac59e5a
abetlen
authored
ac59e5a
View commit details
Copy full SHA for ac59e5a
View code at this point
Browse repository at this point
Commits on Mar 24, 2026
chore: bump version (#2157)
abetlen
authored
d6f46a5
View commit details
Copy full SHA for d6f46a5
View code at this point
Browse repository at this point
feat: expose attention_type parameter in Llama.__init__ (#2143)
Show description for 7b38c31
3 people
authored
7b38c31
View commit details
Copy full SHA for 7b38c31
View code at this point
Browse repository at this point
Commits on Mar 23, 2026
chore: Bump version (#2153)
abetlen
authored
a6b1807
View commit details
Copy full SHA for a6b1807
View code at this point
Browse repository at this point
fix: Qwen 3.5 support (#2152)
Show description for 11e7a55
abetlen
authored
11e7a55
View commit details
Copy full SHA for 11e7a55
View code at this point
Browse repository at this point
feat: Update llama.cpp to ggerganov/llama.cpp@49bfddeca18e62fa3d39114a23e9fcbdf8a22388 (#2151)
Show description for 18aa31e
abetlen
authored
18aa31e
View commit details
Copy full SHA for 18aa31e
View code at this point
Browse repository at this point
Commits on Mar 22, 2026
misc: Add Ruff formatting (#2148)
Show description for a9b4a06
abetlen
authored
a9b4a06
View commit details
Copy full SHA for a9b4a06
View code at this point
Browse repository at this point
Commits on Aug 15, 2025
chore: Bump version
abetlen
committed
c37132b
View commit details
Copy full SHA for c37132b
View code at this point
Browse repository at this point
Commits on Aug 7, 2025
chore: Bump version
abetlen
committed
dfc9bf5
View commit details
Copy full SHA for dfc9bf5
View code at this point
Browse repository at this point
fix: rename op_offloat to op_offload in llama.py (#2046)
sergey21000
authored
30ddd56
View commit details
Copy full SHA for 30ddd56
View code at this point
Browse repository at this point
feat: Add gpt-oss chat format support through strftime_now in chat format by @iamlemec
abetlen
committed
af63792
View commit details
Copy full SHA for af63792
View code at this point
Browse repository at this point
Previous
Next
You can’t perform that action at this time.