deploy¶
- TabularFoundationModel.deploy(instance_type: str | None = None, endpoint_name: str | None = None, hyperparameters: dict[str, Any] | None = None, framework_version: str = '1.6', custom_image_uri: str | None = None, wait: bool = True, inference_mode: Literal['realtime'] = 'realtime', inference_config: dict[str, Any] | None = None, **backend_kwargs) TabularEndpoint[source]¶
Deploy the tabular foundation model to an inference endpoint.
The returned endpoint accepts both labeled
train_dataand the rows to predict. It fits a request-scopedTabularPredictorbefore producing predictions.Only real-time inference is supported. Tabular foundation models such as Mitra require a provisioned instance and cannot be deployed with SageMaker Serverless Inference.
- Parameters:
instance_type (str | None, default = None) – Instance type for the endpoint. Defaults to the model registry value.
endpoint_name (str | None, default = None) – Custom endpoint name. If None, will auto-generate a unique name.
hyperparameters (dict[str, Any] | None, default = None) – Model hyperparameters for inference. Overrides values passed to the constructor.
framework_version (str, default = "1.6") – AutoGluon version, e.g. “1.6”. Uses the official AutoGluon DLC image for this version.
custom_image_uri (str | None, default = None) – Custom Docker image URI for the inference container.
wait (bool, default = True) – Whether to block until the endpoint is ready.
inference_mode (Literal["realtime"], default = "realtime") – Endpoint type. Only
"realtime"is supported.inference_config (dict[str, Any] | None, default = None) – Not supported; must be None.
**backend_kwargs (Any) –
Additional SageMaker arguments:
initial_instance_count: Number of instances for the endpoint. Defaults to 1.volume_size: Size in GB of the EBS volume to use for the endpoint. Ignored for GPU instances.backend_overrides: raw SageMaker request fields for settings without a dedicated argument.Keys: request names from the SageMaker API section below.
Values: request fields in PascalCase, as in the SageMaker API and boto3. Deep-merged over the request built by AutoGluon-Cloud; lists and other non-dict values replace the generated ones.
Example:
{"ProductionVariant": {"ModelDataDownloadTimeoutInSeconds": 1200}}
SageMaker API
CreateModel: registers the model artifact and inference image as a SageMaker model.
CreateEndpointConfig: defines the endpoint’s single ProductionVariant: instance type and count.
CreateEndpoint: launches the endpoint.
The endpoint is billed until
TabularEndpoint.delete_endpoint()deletes it.