deploy¶
- TimeSeriesFoundationModel.deploy(instance_type: str | None = None, endpoint_name: str | None = None, hyperparameters: Dict[str, Any] | None = None, framework_version: str = 'latest', image_uri: str | None = None, wait: bool = True, inference_mode: Literal['realtime', 'serverless'] = 'realtime', inference_config: Dict[str, Any] | None = None, environment: Dict[str, str] | None = None, backend_overrides: Dict[str, Dict[str, Any]] | None = None, **kwargs) TimeSeriesEndpoint[source]¶
Deploy model to an inference endpoint.
- Parameters:
instance_type – Instance type for the endpoint. Defaults to the model registry value. Must be
Nonewheninference_mode="serverless".endpoint_name – Custom endpoint name. If None, will auto-generate a unique name.
hyperparameters – Model hyperparameters for inference. Overrides values passed to the constructor.
framework_version – Container framework version. If ‘latest’, uses the most recent available.
image_uri – Custom Docker image URI for the inference container.
wait – Whether to block until the endpoint is ready.
inference_mode – Endpoint type.
"serverless"provisions a SageMaker Serverless Inference endpoint (no instance management, scales to zero).inference_config – Serverless settings (
memory_size_in_mb,max_concurrency,provisioned_concurrency).environment – Environment variables set in the inference container.
backend_overrides – Raw SageMaker request fields deep-merged over the requests built by AutoGluon-Cloud. Valid keys:
"create_model","production_variant","create_endpoint_config","create_endpoint". Seeautogluon.cloud.TabularCloudPredictor.deploy().**kwargs – Additional deployment arguments (
initial_instance_count,volume_size).