Code: 53968425
Large language models are only as effective as the systems that serve them. As LLM applications move from prototypes to production, engineers face a different class of challenges: rising inference costs, unpredictable latency, lim ... more
English
20.80 €
RRP: 22.35 €
You save 1.54 €

You get 51 loyalty points
Book synopsis
Large language models are only as effective as the systems that serve them. As LLM applications move from prototypes to production, engineers face a different class of challenges: rising inference costs, unpredictable latency, limited GPU memory, retrieval failures, inefficient batching, and performance bottlenecks that cannot be solved by simply adding more hardware.
Inference Engineering & Optimization with LLMs examines the engineering principles behind faster, more efficient, and more scalable language model applications. It brings together retrieval-augmented generation, vector search, embeddings, prompt and context optimization, quantization, KV cache management, multi-GPU parallelism, batching, scheduling, and high-performance model serving into one practical framework.
This handbook is designed around engineering trade-offs rather than one-size-fits-all optimization recipes. It explains how to characterize workloads, identify actual bottlenecks, measure performance, evaluate quality, and make informed decisions about latency, throughput, memory utilization, cost, and model quality.
The book covers established approaches and technologies used across the LLM inference ecosystem, including RAG architectures, vector databases, LangChain, LlamaIndex, vLLM, structured generation, quantization techniques, PagedAttention, tensor parallelism, pipeline parallelism, continuous batching, speculative decoding, and inference benchmarking.
The journey begins with the foundations of inference engineering, including tokenization, prefill, decode, KV caching, time to first token, inter-token latency, throughput, GPU utilization, and workload profiling.
From there, the book moves into retrieval-augmented generation and retrieval infrastructure, explaining chunking, embeddings, dense and sparse retrieval, hybrid search, reranking, vector database architectures, index management, and retrieval-quality evaluation.
The optimization layer then explores prompt and context management, structured outputs, caching, model quantization, memory optimization, KV cache allocation, multi-GPU parallelism, continuous batching, request scheduling, and speculative decoding.
The final stage focuses on sustaining performance through benchmarking, regression testing, production A/B testing, and systematic evaluation of optimization changes.
WHAT'S INSIDE:
Inference lifecycle and performance fundamentals
Retrieval-augmented generation architecture
Vector databases and retrieval infrastructure
Embedding models, chunking, hybrid search, and reranking
LLM orchestration and tool-calling workflows
Prompt compression and context-window optimization
Quantization and precision trade-offs
KV cache and GPU memory optimization
Tensor, pipeline, and data parallelism
Continuous batching and request scheduling
Speculative decoding and throughput engineering
Benchmarking, regression testing, and production evaluation
If you are building, operating, or optimizing production LLM systems, this handbook provides a structured way to move beyond trial-and-error tuning. Learn to profile before optimizing, understand the interactions between application, model, and infrastructure layers, and make performance decisions based on measurable workload requirements.
Build a deeper understanding of LLM inference engineering and turn complex optimization challenges into disciplined, measurable engineering decisions.
Book details
20.80 €
EnglishOsobný odber Bratislava a 13244 dalších
Copyright ©2008-26 najlacnejsie-knihy.sk Všetky práva vyhradenéSúkromieCookies
25 miliónov titulov
Vrátenie do mesiaca
02/210 210 99 (8-15.30h)Nákupný košík ( prázdny )
Nachádzate sa: