Inference Engineering & Optimization with LLMs / Najlacnejšie knihy
Inference Engineering & Optimization with LLMs

Code: 53968425

Inference Engineering & Optimization with LLMs

by Xander N. Lindholm

Large language models are only as effective as the systems that serve them. As LLM applications move from prototypes to production, engineers face a different class of challenges: rising inference costs, unpredictable latency, lim ... more

20.80

RRP: 22.35 €

You save 1.54 €


In stock at our supplier
Shipping in 14 - 21 days
Add to wishlist

You might also like

Give this book as a present today
  1. Order book and choose Gift Order.
  2. We will send you book gift voucher at once. You can give it out to anyone.
  3. Book will be send to donee, nothing more to care about.

Book gift voucher sampleRead more

More about Inference Engineering & Optimization with LLMs

You get 51 loyalty points

Book synopsis

Large language models are only as effective as the systems that serve them. As LLM applications move from prototypes to production, engineers face a different class of challenges: rising inference costs, unpredictable latency, limited GPU memory, retrieval failures, inefficient batching, and performance bottlenecks that cannot be solved by simply adding more hardware.

Inference Engineering & Optimization with LLMs examines the engineering principles behind faster, more efficient, and more scalable language model applications. It brings together retrieval-augmented generation, vector search, embeddings, prompt and context optimization, quantization, KV cache management, multi-GPU parallelism, batching, scheduling, and high-performance model serving into one practical framework.

This handbook is designed around engineering trade-offs rather than one-size-fits-all optimization recipes. It explains how to characterize workloads, identify actual bottlenecks, measure performance, evaluate quality, and make informed decisions about latency, throughput, memory utilization, cost, and model quality.

The book covers established approaches and technologies used across the LLM inference ecosystem, including RAG architectures, vector databases, LangChain, LlamaIndex, vLLM, structured generation, quantization techniques, PagedAttention, tensor parallelism, pipeline parallelism, continuous batching, speculative decoding, and inference benchmarking.

The journey begins with the foundations of inference engineering, including tokenization, prefill, decode, KV caching, time to first token, inter-token latency, throughput, GPU utilization, and workload profiling.

From there, the book moves into retrieval-augmented generation and retrieval infrastructure, explaining chunking, embeddings, dense and sparse retrieval, hybrid search, reranking, vector database architectures, index management, and retrieval-quality evaluation.

The optimization layer then explores prompt and context management, structured outputs, caching, model quantization, memory optimization, KV cache allocation, multi-GPU parallelism, continuous batching, request scheduling, and speculative decoding.

The final stage focuses on sustaining performance through benchmarking, regression testing, production A/B testing, and systematic evaluation of optimization changes.

WHAT'S INSIDE:

Inference lifecycle and performance fundamentals

Retrieval-augmented generation architecture

Vector databases and retrieval infrastructure

Embedding models, chunking, hybrid search, and reranking

LLM orchestration and tool-calling workflows

Prompt compression and context-window optimization

Quantization and precision trade-offs

KV cache and GPU memory optimization

Tensor, pipeline, and data parallelism

Continuous batching and request scheduling

Speculative decoding and throughput engineering

Benchmarking, regression testing, and production evaluation

If you are building, operating, or optimizing production LLM systems, this handbook provides a structured way to move beyond trial-and-error tuning. Learn to profile before optimizing, understand the interactions between application, model, and infrastructure layers, and make performance decisions based on measurable workload requirements.

Build a deeper understanding of LLM inference engineering and turn complex optimization challenges into disciplined, measurable engineering decisions.

Book details

20.80



Osobný odber Bratislava a 13244 dalších

Copyright ©2008-26 najlacnejsie-knihy.sk Všetky práva vyhradenéSúkromieCookies


Môj účet: Prihlásiť sa
Všetky knihy sveta na jednom mieste. Navyše za skvelé ceny.

Nákupný košík ( prázdny )

Vyzdvihnutie v Zásielkovni
zadarmo nad 59,99 €.

Nachádzate sa: