← AWS Certified Generative AI Developer - Professional
AWS · objective · 12% of the exam
Operational Efficiency and Optimization for GenAI Applications — AWS Certified Generative AI Developer - Professional
The official AWS documentation our Operational Efficiency and Optimization for GenAI Applications practice questions are cited to. Review the primary sources, then practise.
Official references for this objective
-
Amazon Web Services — Prompt caching for faster model inference - Amazon Bedrock
explicit — Disables the automatic breakpoint. Only explicit breakpoints are used for cache reads and writes.
-
Amazon Web Services — Effectively use prompt caching on Amazon Bedrock | Artificial Intelligence
At times of high demand, these optimizations may lead to increased cache writes
-
Amazon Web Services — Amazon Bedrock Cost Optimization - AWS
Intelligent Prompt Routing can reduce costs by up to 30% without compromising on accuracy.
-
Amazon Web Services — Autoscale an asynchronous endpoint - Amazon SageMaker AI
your endpoint only initiates scaling up from zero after the number of backlog requests exceeds the target tracking value
-
Amazon Web Services — Customize a model with distillation in Amazon Bedrock - Amazon Bedrock
the teacher model specified in your model distillation job must match the model used in the invocation log. If they don't match, the invocation logs
-
Amazon Web Services — Deploy models for real-time inference - Amazon SageMaker AI
you can optimize resource utilization by tailoring how the required CPU cores, accelerators, and memory are allocated to the model
-
Amazon Web Services — Increase model invocation capacity with Provisioned Throughput in Amazon Bedrock - Amazon Bedrock
You're billed hourly for a Provisioned Throughput that you purchase.