[pull] main from abetlen:main by pull[bot] · Pull Request #3 · imotai/llama-cpp-python

pull · 2023-11-03T02:27:53Z

See Commits and Changes for more details.

Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

Bumps [pypa/cibuildwheel](https://github.com/pypa/cibuildwheel) from 2.18.1 to 2.19.1. - [Release notes](https://github.com/pypa/cibuildwheel/releases) - [Changelog](https://github.com/pypa/cibuildwheel/blob/main/docs/changelog.md) - [Commits](pypa/cibuildwheel@v2.18.1...v2.19.1) --- updated-dependencies: - dependency-name: pypa/cibuildwheel dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <[email protected]> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

…be enabled by default for embedding models.

…to main

fix: :( fix fix fix: Copy runtime dlls on windows fix: Add explicit copy command for windows dlls fix fix fix: >:( fix fix fix fix Update path on windows check dll dependancies fix: Update PATH on win32 ci: Update test.yaml

…as released before request completes causing SEGFAULT

@iamlemec

…rmat by @iamlemec

* Fix model download in test workflow * Use hf CLI in test workflow * Use hf CLI name in CI and docs * Reference PR in changelog

#2150) * fix(ci): use supported macos runner label * fix(ci): add apple silicon macos test coverage * fix(ci): run standard macos tests on apple silicon * fix(ci): simplify apple silicon macos install * fix(ci): disable ggml native on apple silicon runner * docs: update changelog for macos ci runner fix

* Add Ruff formatting and safe lint baseline * Update changelog for Ruff setup

* Update llama.cpp and sync bindings * Clean up binding compatibility shims * Remove flash attention property shim * Remove mtmd verbosity shim * Add docstrings for new bindings * Format Ruff files and add changelog entry

* ci: add riscv64 wheel builds to release workflow Add a build_wheels_riscv64 job mirroring the existing arm64 QEMU-based build. Uses cibuildwheel with QEMU emulation for linux/riscv64, targeting CPython 3.10-3.14 on manylinux. Closes #2138 * ci: use cibuildwheel 3.1.2 for riscv64 wheels * docs: update changelog for riscv64 wheel PR --------- Co-authored-by: abetlen <[email protected]>

* fix: handle Qwen 3.5 hybrid prefix reuse * test: fix Qwen runtime unit mocks * test: drop Qwen runtime unit tests * docs: credit Qwen fix contributors in changelog * docs/tests: update default Qwen model to 3.5 0.8B * test: rebaseline Qwen 3.5 outputs * test: stabilize low-level Qwen sampling check * test: tighten Qwen 3.5 completion prompts

* fix(ci): harden release wheel workflow * fix(ci): document and pin release wheel baselines * fix(ci): speed up release arch builds * fix(ci): split riscv64 by python version * fix(ci): sanitize riscv64 artifact names

* fix(ci): harden cuda wheel workflow * fix(ci): pin cuda toolkit versions accurately * fix(ci): resolve exact cuda toolkit installs * fix(ci): align cuda toolkit roots and tags * fix(ci): pin cuda packages to nvidia label * fix(ci): allow cuda solver to mix non-cuda deps

* fix(ci): harden docker build workflow * docs: update changelog for ci workflows

* feat: expose attention_type parameter in Llama.__init__ * docs: preserve attention_type in pickled state * docs: update changelog for attention_type --------- Co-authored-by: Victor Biederbeck <[email protected]> Co-authored-by: abetlen <[email protected]>

…ent arches and one PTX target for forward compatibility (#2158) * fix(ci): shrink CUDA wheel fatbins * docs: update changelog for cuda wheel size fix

* Fix embedding models without KV memory * Add changelog entry for embedding memory fix

* Update llama.cpp to c0159f9c1 * Add changelog entry for llama.cpp update

* fix(ci): publish distinct manylinux and musllinux cpu wheels * docs: add changelog entry for linux wheel repair fix

* ci: publish CPU wheels as py3-none * docs: add changelog entry for py3-none wheel tags

* refactor: replace deprecated llama.cpp references * docs: update changelog for recent llama.cpp changes

* feat: Update llama.cpp to ggml-org/llama.cpp@3bd9aa1f9 * docs: Update changelog for llama.cpp bump

pull bot added ⤵️ pull merge-conflict Resolve conflicts manually labels Nov 4, 2023

abetlen force-pushed the main branch 5 times, most recently from 4408d7a to cc0fe43 Compare November 14, 2023 20:30

abetlen force-pushed the main branch from 0188482 to c96b2da Compare April 17, 2024 14:06

dependabot bot and others added 11 commits June 21, 2024 12:10

feat: Update llama_cpp.py bindings

04959f1

feat: Update llama.cpp

117cbb2

fix(server): Update embeddings=False by default. Embeddings should …

bf5e0bb

…be enabled by default for embedding models.

fix(ci): Fix the CUDA workflow (#1551)

73ddf29

misc: Install shared libraries to lib subdirectory

c546c94

Merge branch 'main' of https://github.com/abetlen/llama-cpp-python in…

92bad6e

…to main

fix: Update shared library rpath

139774b

fix: force $ORIGIN rpath for shared library files

d5f6a15

fix: Fix installation location for shared libraries

e51f200

fix: Fix RPATH so it works on macos

73fe013

abetlen force-pushed the main branch from 5a7ad37 to 73fe013 Compare July 2, 2024 05:33

fix: Copy dependencies for windows

dc20e8c

fix: :( fix fix fix: Copy runtime dlls on windows fix: Add explicit copy command for windows dlls fix fix fix: >:( fix fix fix fix Update path on windows check dll dependancies fix: Update PATH on win32 ci: Update test.yaml

abetlen force-pushed the main branch from 5a7ad37 to dc20e8c Compare July 2, 2024 05:39

abetlen added 8 commits July 2, 2024 02:49

fix(server): Fix bug in FastAPI streaming response where dependency w…

296304b

…as released before request completes causing SEGFAULT

feat: Update llama.cpp

bd5d17b

chore: Bump version

b4cc923

fix(ci): Use LLAMA_CUDA for cuda wheels

4fb6fc1

fix(misc): Fix type errors

387d01d

feat: Update llama.cpp

8992a1a

fix(ci): Update macos image (macos-11 is removed)

3a551eb

chore: Bump version

01bddd6

abetlen and others added 30 commits August 7, 2025 06:42

misc: Update pypi downloads badge

d12ca47

misc: Add Python 3.13 classifier tag

68e89e8

feat: Add gpt-oss chat format support through strftime_now in chat fo…

af63792

…rmat by @iamlemec

fix: rename op_offloat to op_offload in llama.py (#2046)

30ddd56

chore: Bump version

dfc9bf5

feat: Update llama.cpp

ce6fd8b

chore: Bump version

c37132b

fix(ci): Rename huggingface-cli to hf (#2149)

ca3b00a

* Fix model download in test workflow * Use hf CLI in test workflow * Use hf CLI name in CI and docs * Reference PR in changelog

misc: Add Ruff formatting (#2148)

a9b4a06

* Add Ruff formatting and safe lint baseline * Update changelog for Ruff setup

feat: Update llama.cpp to ggml-org/llama.cpp@49bfdde (#2151)

18aa31e

* Update llama.cpp and sync bindings * Clean up binding compatibility shims * Remove flash attention property shim * Remove mtmd verbosity shim * Add docstrings for new bindings * Format Ruff files and add changelog entry

chore: Bump version (#2153)

a6b1807

fix(ci): release wheel workflow (#2154)

f0391c5

* fix(ci): harden release wheel workflow * fix(ci): document and pin release wheel baselines * fix(ci): speed up release arch builds * fix(ci): split riscv64 by python version * fix(ci): sanitize riscv64 artifact names

fix(ci): docker build workflow (#2156)

ccc6bc0

* fix(ci): harden docker build workflow * docs: update changelog for ci workflows

chore: bump version (#2157)

d6f46a5

fix(ci): reduce CUDA binary wheel size only including cubins for curr…

5f9c231

…ent arches and one PTX target for forward compatibility (#2158) * fix(ci): shrink CUDA wheel fatbins * docs: update changelog for cuda wheel size fix

fix: handle embedding models without KV memory (#2160)

ac59e5a

* Fix embedding models without KV memory * Add changelog entry for embedding memory fix

feat: Update llama.cpp to ggml-org/llama.cpp@c0159f9 (#2161)

c670222

* Update llama.cpp to c0159f9c1 * Add changelog entry for llama.cpp update

Bump version to 0.3.19 (#2162)

f54421b

fix(ci): publish distinct manylinux and musllinux cpu wheels (#2165)

fcd932a

* fix(ci): publish distinct manylinux and musllinux cpu wheels * docs: add changelog entry for linux wheel repair fix

ci: publish release wheels as py3-none (#2166)

7613aca

* ci: publish CPU wheels as py3-none * docs: add changelog entry for py3-none wheel tags

feat(server): add model-load chat_template_kwargs (#2168)

7257ba9

feat: Update llama.cpp to ggml-org/llama.cpp@f49e917 (#2169)

100b275

fix(misc): replace deprecated llama.cpp references (#2170)

08e088c

* refactor: replace deprecated llama.cpp references * docs: update changelog for recent llama.cpp changes

chore: bump version to 0.3.20 (#2171)

02d6bee

feat: Update llama.cpp to ggml-org/llama.cpp@3bd9aa1f9 (#2176)

1bcc5bc

* feat: Update llama.cpp to ggml-org/llama.cpp@3bd9aa1f9 * docs: Update changelog for llama.cpp bump

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[pull] main from abetlen:main#3

[pull] main from abetlen:main#3
pull[bot] wants to merge 910 commits intoimotai:mainfrom
abetlen:main

pull bot commented Nov 3, 2023 •

edited

Loading

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants

Conversation

pull bot commented Nov 3, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants

pull bot commented Nov 3, 2023 •

edited

Loading