fit

TabularCloudPredictor.fit(train_data: str | Path | DataFrame | None = None, *, tuning_data: str | Path | DataFrame | None = None, predictor_init_args: Dict[str, Any], predictor_fit_args: Dict[str, Any] | None = None, image_column: str | None = None, leaderboard: bool = True, framework_version: str = 'latest', job_name: str | None = None, instance_type: str = 'ml.m5.2xlarge', instance_count: int | str = 'auto', volume_size: int = 256, image_uri: str | None = None, timeout: int = 86400, wait: bool = True, environment: Dict[str, str] | None = None, use_spot_instances: bool = False, max_wait: int | None = None, sagemaker_overrides: Dict[str, Dict[str, Any]] | None = None, **kwargs) → CloudPredictor

Fit the predictor in a SageMaker training job.

Parameters:
  • train_data (Union[str, pathlib.Path, pd.DataFrame]) – Training data, as a DataFrame or local/S3 path to a data file.

  • tuning_data (Optional[Union[str, pathlib.Path, pd.DataFrame]], default = None) – Optional tuning data.

  • predictor_init_args (dict) – Init args for the predictor.

  • predictor_fit_args (Optional[dict], default = None) – Additional fit args forwarded to the underlying predictor’s fit(). Must NOT contain train_data or tuning_data — pass those as explicit arguments above.

  • leaderboard (bool, default = True) – Whether to include the leaderboard in the output artifact

  • framework_version (str, default = latest) – Training container version of autogluon. If latest, will use the latest available container version. If provided a specific version, will use this version. If image_uri is set, this argument will be ignored.

  • job_name (str, default = None) – Name of the launched training job. If None, CloudPredictor creates one with a predictor-specific prefix.

  • instance_type (str, default = 'ml.m5.2xlarge') – Instance type the predictor will be trained on with SageMaker.

  • instance_count (Union[int, str], default = "auto") – Number of instances used to fit the predictor. If “auto”, the backend decides the instance count.

  • volume_size (int, default = 256) – Size in GB of the EBS volume to use for storing input data during training. Must be large enough to store training data if File Mode is used (which is the default).

  • image_uri (Optional[str], default = None) – Custom training container image. If None, the official AutoGluon DLC for framework_version is used.

  • timeout (int, default = 24*60*60) – Timeout in seconds for training. This timeout doesn’t include time for pre-processing or launching up the training job.

  • wait (bool, default = True) – Whether the call should wait until the job completes To be noticed, the function won’t return immediately because there are some preparations needed prior fit. Use get_fit_job_status to get job status.

  • environment (Optional[Dict[str, str]], default = None) – Environment variables set in the training container.

  • use_spot_instances (bool, default = False) – Whether to train on managed spot instances.

  • max_wait (Optional[int], default = None) – Maximum seconds to wait for spot capacity plus training time. Defaults to timeout. Requires use_spot_instances=True.

  • sagemaker_overrides (Optional[Dict[str, Dict[str, Any]]], default = None) – Escape hatch for SageMaker settings without a dedicated argument. Maps "create_training_job" to raw CreateTrainingJob request fields in snake_case (as in sagemaker.core.shapes), which are deep-merged over the request built by AutoGluon-Cloud, e.g. {"create_training_job": {"retry_strategy": {"maximum_retry_attempts": 2}}}.

Return type:

CloudPredictor object. Returns self.