deploy

TimeSeriesFoundationModel.deploy(instance_type: str | None = None, endpoint_name: str | None = None, hyperparameters: Dict[str, Any] | None = None, framework_version: str = 'latest', image_uri: str | None = None, wait: bool = True, inference_mode: Literal['realtime', 'serverless'] = 'realtime', inference_config: Dict[str, Any] | None = None, environment: Dict[str, str] | None = None, sagemaker_overrides: Dict[str, Dict[str, Any]] | None = None, **kwargs) → TimeSeriesEndpoint[source]

Deploy model to an inference endpoint.

Parameters:
  • instance_type – Instance type for the endpoint. Defaults to the model registry value. Must be None when inference_mode="serverless".

  • endpoint_name – Custom endpoint name. If None, will auto-generate a unique name.

  • hyperparameters – Model hyperparameters for inference. Overrides values passed to the constructor.

  • framework_version – Container framework version. If ‘latest’, uses the most recent available.

  • image_uri – Custom Docker image URI for the inference container.

  • wait – Whether to block until the endpoint is ready.

  • inference_mode – Endpoint type. "serverless" provisions a SageMaker Serverless Inference endpoint (no instance management, scales to zero).

  • inference_config – Serverless settings (memory_size_in_mb, max_concurrency, provisioned_concurrency).

  • environment – Environment variables set in the inference container.

  • sagemaker_overrides – Raw SageMaker request fields deep-merged over the requests built by AutoGluon-Cloud. Valid keys: "create_model", "production_variant", "create_endpoint_config", "create_endpoint". See autogluon.cloud.TabularCloudPredictor.deploy().

  • **kwargs – Additional deployment arguments (initial_instance_count, volume_size).