LIBRISTO
LIBROAMANTO
obligatorisch
Werden Sie Teil einer Gemeinschaft von Buchliebhabern aus der ganzen Welt und erhalten Sie eine Reihe von Vorteilen. Konto kostenlos anlegen
0
Österreichische Post 5.49 GLS-Kurier 4.99 DPD-Kurier 4.49 DPD-Stelle 3.49

Inference Engineering & Optimization with LLMs

A Handbook to Retrieval-Augmented Generation, Quantization, Parallelism, and High-Performance Model Serving.

Sprache EnglischEnglisch
Buch Broschur
Verlag Independently published, September 2026
Large language models are only as effective as the systems that serve them. As LLM applications move... Vollständige Beschreibung
? points 51 b Neu Neu
20.69 inkl. MwSt.
Externes Lager Wir versenden in 14-21 Tagen

Bis zu 30 Tage Rückgaberecht

Large language models are only as effective as the systems that serve them. As LLM applications move from prototypes to production, engineers face a different class of challenges: rising inference costs, unpredictable latency, limited GPU memory, retrieval failures, inefficient batching, and performance bottlenecks that cannot be solved by simply adding more hardware.

Inference Engineering & Optimization with LLMs examines the engineering principles behind faster, more efficient, and more scalable language model applications. It brings together retrieval-augmented generation, vector search, embeddings, prompt and context optimization, quantization, KV cache management, multi-GPU parallelism, batching, scheduling, and high-performance model serving into one practical framework.

This handbook is designed around engineering trade-offs rather than one-size-fits-all optimization recipes. It explains how to characterize workloads, identify actual bottlenecks, measure performance, evaluate quality, and make informed decisions about latency, throughput, memory utilization, cost, and model quality.

The book covers established approaches and technologies used across the LLM inference ecosystem, including RAG architectures, vector databases, LangChain, LlamaIndex, vLLM, structured generation, quantization techniques, PagedAttention, tensor parallelism, pipeline parallelism, continuous batching, speculative decoding, and inference benchmarking.

The journey begins with the foundations of inference engineering, including tokenization, prefill, decode, KV caching, time to first token, inter-token latency, throughput, GPU utilization, and workload profiling.

From there, the book moves into retrieval-augmented generation and retrieval infrastructure, explaining chunking, embeddings, dense and sparse retrieval, hybrid search, reranking, vector database architectures, index management, and retrieval-quality evaluation.

The optimization layer then explores prompt and context management, structured outputs, caching, model quantization, memory optimization, KV cache allocation, multi-GPU parallelism, continuous batching, request scheduling, and speculative decoding.

The final stage focuses on sustaining performance through benchmarking, regression testing, production A/B testing, and systematic evaluation of optimization changes.

WHAT'S INSIDE:

Inference lifecycle and performance fundamentals

Retrieval-augmented generation architecture

Vector databases and retrieval infrastructure

Embedding models, chunking, hybrid search, and reranking

LLM orchestration and tool-calling workflows

Prompt compression and context-window optimization

Quantization and precision trade-offs

KV cache and GPU memory optimization

Tensor, pipeline, and data parallelism

Continuous batching and request scheduling

Speculative decoding and throughput engineering

Benchmarking, regression testing, and production evaluation

If you are building, operating, or optimizing production LLM systems, this handbook provides a structured way to move beyond trial-and-error tuning. Learn to profile before optimizing, understand the interactions between application, model, and infrastructure layers, and make performance decisions based on measurable workload requirements.

Build a deeper understanding of LLM inference engineering and turn complex optimization challenges into disciplined, measurable engineering decisions.

Schauspielerin & Polyglotte
EWA KASP für
Video abspielen
Ewa Kasp
Libristo bietet die größte Auswahl an fremdsprachiger Literatur an. Deshalb kaufe ich meine Bücher hier ein.

Informationen zum Buch

Vollständiger Name Inference Engineering & Optimization with LLMs
Sprache Englisch
Einband Buch - Broschur
Datum der Veröffentlichung 2026
Anzahl der Seiten 276
EAN 9798175305259
Libristo-Code 53968425
Gewicht 374
Abmessungen 152 x 229 x 15
Verschenken Sie dieses Buch noch heute
Es ist ganz einfach
1 Legen Sie das Buch in Ihren Warenkorb und wählen Sie den Versand als Geschenk 2 Wir schicken Ihnen umgehend einen Gutschein 3 Das Buch wird an die Adresse des beschenkten Empfängers geliefert

Anmeldung

Melden Sie sich bei Ihrem Konto an. Sie haben noch kein Libristo-Konto? Erstellen Sie es jetzt!

 
obligatorisch
obligatorisch

Sie haben kein Konto? Nutzen Sie die Vorteile eines Libristo-Kontos!

Mit einem Libristo-Konto haben Sie alles unter Kontrolle.

Erstellen Sie ein Libristo-Konto
Buchberater Libroamiko
Hallo, ich bin Libroamiko, kann ich helfen?