Transformers falcon text-generation-inference

Falcon 40B Base Model GGUF

These files are GGUF format quantized model files for TII's tiiuae/Falcon 40B base model.

<!-- README_GGUF.md-about-gguf start -->

About GGUF

GGUF is a new format introduced by the llama.cpp team on August 21st 2023. It is a replacement for GGML, which is no longer supported by llama.cpp.

The key benefit of GGUF is that it is a extensible, future-proof format which stores more information about the model as metadata. It also includes significantly improved tokenization code, including for the first time full support for special tokens. This should improve performance, especially with models that use new special tokens and implement custom prompt templates.

As of August 25th, here is a list of clients and libraries that are known to support GGUF:

The clients and libraries below are expecting to add GGUF support shortly: <!-- README_GGUF.md-about-gguf end -->

<!-- repositories-available start -->

Repositories available

<!-- repositories-available end -->