deploy¶
- TabularCloudPredictor.deploy(predictor_path: str | None = None, endpoint_name: str | None = None, framework_version: str | None = None, instance_type: str | None = None, initial_instance_count: int = 1, custom_image_uri: str | None = None, volume_size: int | None = None, wait: bool = True, inference_mode: Literal['realtime', 'serverless'] = 'realtime', inference_config: dict[str, Any] | None = None, backend_overrides: dict[str, dict[str, Any]] | None = None) EndpointT¶
Deploy a predictor to an inference endpoint.
Returns a handle to the endpoint; call its
predict()for low-latency inference, thendelete_endpoint()to tear it down.- Parameters:
predictor_path (str) – Path to the predictor tarball you want to deploy. Path can be both a local path or a S3 location. If None, will deploy the most recent trained predictor trained with fit().
endpoint_name (str) – The endpoint name to use for the deployment. If None, CloudPredictor creates one with a predictor-specific prefix.
framework_version (str, optional) – AutoGluon version, e.g. “1.6”. Inference uses the official AutoGluon DLC image for this version. Defaults to the version used by fit(). If custom_image_uri is set, this argument will be ignored.
instance_type (str | None, default = None) – Instance to be deployed for the endpoint. Defaults to
ml.m5.2xlarge. Must beNonewheninference_mode="serverless".initial_instance_count (int, default = 1,) – Initial number of instances to be deployed for the endpoint. Ignored when
inference_mode="serverless".custom_image_uri (str | None, default = None,) – Custom image to use to deploy endpoint with. If not specified, with use official DLC image: https://github.com/aws/deep-learning-containers/blob/master/available_images.md#autogluon-inference-containers
volume_size (int, default = None) – Size in GB of the EBS volume to use for the endpoint. SageMaker GPU instance endpoint currently doesn’t support specifying volume_size. Will ignore in such cases.
wait (Bool, default = True,) – Whether to wait for the endpoint to be deployed. To be noticed, the function won’t return immediately because there are some preparations needed prior deployment.
inference_mode ({"realtime", "serverless"}, default = "realtime") – Endpoint type.
"serverless"provisions a SageMaker Serverless Inference endpoint (no instance management, scales to zero).inference_config (dict[str, Any] | None, default = None) – Serverless settings (
memory_size_in_mb,max_concurrency,provisioned_concurrency).backend_overrides (dict[str, dict[str, Any]] | None, default = None) –
Raw SageMaker request fields for settings without a dedicated argument.
Keys: request names from the SageMaker API section below.
Values: request fields in PascalCase, as in the SageMaker API and boto3. Deep-merged over the request built by AutoGluon-Cloud; lists and other non-dict values replace the generated ones.
Example:
{"ProductionVariant": {"ModelDataDownloadTimeoutInSeconds": 1200}}
- Returns:
TabularEndpoint | TimeSeriesEndpoint | MultiModalEndpoint – Handle to the deployed endpoint, matching the predictor type.
SageMaker API
CreateModel: registers the model artifact and inference image as a SageMaker model.
CreateEndpointConfig: defines the endpoint’s single ProductionVariant: instance type and count, or the serverless settings.
CreateEndpoint: launches the endpoint.
The endpoint is billed until
delete_endpoint()of the returned endpoint deletes it.