Kjarni 0.1.3
dotnet add package Kjarni --version 0.1.3
NuGet\Install-Package Kjarni -Version 0.1.3
<PackageReference Include="Kjarni" Version="0.1.3" />
<PackageVersion Include="Kjarni" Version="0.1.3" />
<PackageReference Include="Kjarni" />
paket add Kjarni --version 0.1.3
#r "nuget: Kjarni, 0.1.3"
#:package Kjarni@0.1.3
#addin nuget:?package=Kjarni&version=0.1.3
#tool nuget:?package=Kjarni&version=0.1.3
Kjarni
Local AI inference for .NET. No Python, no ONNX Runtime, no CUDA, no cloud.
Embeddings, classification, semantic search, reranking, and local LLM chat — running in your process against a native library. The only runtime dependency is glibc.
dotnet add package Kjarni
Native binaries for linux-x64, linux-arm64, win-x64 and osx-arm64 ship inside the package
and are selected automatically. There is no second install step, no model server to run, and no
API key.
Quick start
using Kjarni;
using var classifier = new Classifier("roberta-sentiment");
Console.WriteLine(classifier.Classify("I love this product!"));
// positive (98.5%)
Models download on first use and are cached locally. After that, nothing touches the network.
Embeddings
using var embedder = new Embedder("minilm-l6-v2");
float[] vector = embedder.Encode("Hello world"); // 384 dimensions
Console.WriteLine(embedder.Similarity("doctor", "physician")); // 0.8598
var docs = new[] { "How do I reset my password?", "What is your refund policy?" };
var vectors = embedder.EncodeBatch(docs);
var query = embedder.Encode("I need to change my login credentials");
var score = Embedder.CosineSimilarity(query, vectors[0]); // 0.5981
Pass whole collections to EncodeBatch rather than looping — Kjarni batches natively.
Microsoft.Extensions.AI and Semantic Kernel
The companion package implements IEmbeddingGenerator<string, Embedding<float>>, so Kjarni drops
into the standard .NET AI abstractions — including Semantic Kernel — with no cloud call and no
Ollama daemon:
dotnet add package Kjarni.Extensions.AI
builder.Services.AddKjarniEmbeddingGenerator("minilm-l6-v2");
See Kjarni.Extensions.AI.
Classification
using var classifier = new Classifier("roberta-sentiment");
Console.WriteLine(classifier.Classify("Terrible quality.").ToJson());
// {"label": "negative", "score": 0.9408, "predictions": [...]}
using var multi = new Classifier("bert-sentiment-multilingual");
Console.WriteLine(multi.Classify("Esta es la peor compra que he hecho."));
// 1 star (94.1%)
using var toxic = new Classifier("toxic-bert");
Console.WriteLine(toxic.Classify("You are an idiot").ToDetailedString());
// toxic 98.61% ███████████████████████████████████████
// insult 96.00% ██████████████████████████████████████
// obscene 75.64% ██████████████████████████████
// severe_toxic 4.56% █
// identity_hate 1.41%
using var emotion = new Classifier("distilroberta-emotion");
Console.WriteLine(emotion.Classify("I just got promoted!"));
// surprise (50.7%)
Chat
Local LLM chat, in-process. No daemon, no API key.
using var chat = new Chat("llama3.2-3b-instruct");
Console.WriteLine(chat.Send("Explain retrieval-augmented generation in one sentence."));
Streaming, token by token:
chat.Stream("Write a haiku about Reykjavík.", token =>
{
Console.Write(token);
return true; // return false to stop generation early
});
Multi-turn conversations keep their own history:
var convo = chat.Conversation();
convo.Send("My name is Ólafur.");
Console.WriteLine(convo.Send("What is my name?")); // remembers
convo.Clear(); // keeps the system prompt
Sampling is configurable via GenerationConfig:
var config = GenerationConfig.Default() with { Temperature = 0.2f, MaxNewTokens = 512 };
chat.Send("Summarise this changelog.", config);
GenerationConfig.Greedy() and GenerationConfig.Creative() are provided as presets.
Reranking
using var reranker = new Reranker();
var results = reranker.Rerank("What is machine learning?", new[] {
"Machine learning is a subset of artificial intelligence.",
"The weather today is sunny.",
});
// 10.5139: Machine learning is a subset of artificial intelligence.
// -11.1001: The weather today is sunny.
Cross-encoder reranking scores query and document together, which is markedly more accurate than comparing embeddings — worth applying to the top ~50 results of a search.
Index and search
using var indexer = new Indexer(model: "minilm-l6-v2", quiet: true);
indexer.Create("my_index", new[] { "docs/" });
using var searcher = new Searcher(
model: "minilm-l6-v2",
rerankerModel: "minilm-l6-v2-cross-encoder");
var results = searcher.Search("my_index", "how do returns work?", mode: SearchMode.Hybrid);
Search modes: Semantic, Keyword (BM25), Hybrid.
Models
Sizes are what the model occupies on disk after download.
| Task | Model | Dimensions | On disk |
|---|---|---|---|
| Embeddings | minilm-l6-v2 |
384 | 88 MB |
| Embeddings | mpnet-base-v2 |
768 | 419 MB |
| Embeddings | nomic-embed-text |
768 | 523 MB |
| Embeddings (multilingual) | bge-m3 |
1024 | ~2 GB |
| Reranking | minilm-l6-v2-cross-encoder |
— | 88 MB |
| Sentiment (binary) | distilbert-sentiment |
— | 257 MB |
| Sentiment (3-class) | roberta-sentiment |
— | 479 MB |
| Sentiment (multilingual) | bert-sentiment-multilingual |
— | 641 MB |
| Emotion (7-class) | distilroberta-emotion |
— | 317 MB |
| Emotion (28-class) | roberta-emotions |
— | 478 MB |
| Toxicity | toxic-bert |
— | 419 MB |
Chat models range from qwen2.5-0.5b-instruct up through llama3.2-3b-instruct, phi3.5-mini,
mistral-7b and deepseek-r1-8b. Run kjarni model list with the CLI for the full registry.
Start with minilm-l6-v2 for embeddings — at 384 dimensions it is fast on CPU and the quality gap
against much larger models is smaller than people expect.
GPU
using var embedder = new Embedder("minilm-l6-v2", device: "gpu");
GPU inference uses WebGPU — Vulkan on Linux, DX12 or Vulkan on Windows, Metal on macOS. CUDA is not required and is not used.
Configuration
using var embedder = new Embedder("minilm-l6-v2", cacheDir: "/my/models");
using var quiet = new Embedder("minilm-l6-v2", quiet: true);
KJARNI_CACHE_DIR overrides the default cache location. HF_TOKEN is used for gated models.
Platform support
| Platform | Shipped | GPU backend |
|---|---|---|
| Linux x64 | Yes | Vulkan |
| Linux arm64 | Yes | Vulkan |
| Windows x64 | Yes | DX12 / Vulkan |
| macOS arm64 | Yes | Metal |
The native library links only against glibc — no CUDA, no BLAS, no ONNX Runtime, no Python.
Links
- Source and issues
- Kjarni.Extensions.AI — Microsoft.Extensions.AI provider
- kjarni.ai
MIT licensed.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net8.0
- No dependencies.
NuGet packages (1)
Showing the top 1 NuGet packages that depend on Kjarni:
| Package | Downloads |
|---|---|
|
Kjarni.Extensions.AI
Microsoft.Extensions.AI provider for Kjarni — local embedding generation for .NET with no Python, no ONNX, and no cloud. Plugs into Semantic Kernel and any IEmbeddingGenerator consumer. |
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.1.3 | 43 | 8/26/2026 |
| 0.1.3-preview.1 | 36 | 8/25/2026 |
| 0.1.0 | 226 | 2/12/2026 |
| 0.1.0-preview.2 | 94 | 2/10/2026 |
| 0.1.0-preview.1 | 81 | 2/7/2026 |