deploy¶
- TimeSeriesFoundationModel.deploy(instance_type: str | None = None, endpoint_name: str | None = None, hyperparameters: dict[str, Any] | None = None, framework_version: str = '1.6', custom_image_uri: str | None = None, wait: bool = True, inference_mode: Literal['realtime', 'serverless'] = 'realtime', inference_config: dict[str, Any] | None = None, **backend_kwargs) TimeSeriesEndpoint[source]¶
Deploy model to an inference endpoint.
- Parameters:
instance_type (str | None, default = None) – Instance type for the endpoint. Defaults to the model registry value. Must be
Nonewheninference_mode="serverless".endpoint_name (str | None, default = None) – Custom endpoint name. If None, will auto-generate a unique name.
hyperparameters (dict[str, Any] | None, default = None) – Model hyperparameters for inference. Overrides values passed to the constructor.
framework_version (str, default = "1.6") – AutoGluon version, e.g. “1.6”. Uses the official AutoGluon DLC image for this version.
custom_image_uri (str | None, default = None) – Custom Docker image URI for the inference container.
wait (bool, default = True) – Whether to block until the endpoint is ready.
inference_mode (Literal["realtime", "serverless"], default = "realtime") – Endpoint type.
"serverless"provisions a SageMaker Serverless Inference endpoint (no instance management, scales to zero).inference_config (dict[str, Any] | None, default = None) – Serverless settings (
memory_size_in_mb,max_concurrency,provisioned_concurrency).**backend_kwargs (Any) –
Additional SageMaker arguments:
initial_instance_count: Number of instances for the endpoint. Defaults to 1. Ignored wheninference_mode="serverless".volume_size: Size in GB of the EBS volume to use for the endpoint. Ignored for GPU instances.backend_overrides: raw SageMaker request fields for settings without a dedicated argument.Keys: request names from the SageMaker API section below.
Values: request fields in PascalCase, as in the SageMaker API and boto3. Deep-merged over the request built by AutoGluon-Cloud; lists and other non-dict values replace the generated ones.
Example:
{"ProductionVariant": {"ModelDataDownloadTimeoutInSeconds": 1200}}
SageMaker API
CreateModel: registers the model artifact and inference image as a SageMaker model.
CreateEndpointConfig: defines the endpoint’s single ProductionVariant: instance type and count, or the serverless settings.
CreateEndpoint: launches the endpoint.
The endpoint is billed until
TimeSeriesEndpoint.delete_endpoint()deletes it.