ホーム > Databricks Machine Learning Professional 試験の主な試験内容は何ですか?-ktest
迅速レーダー
会社案内
ソフト版
購入問題
品質保証
全商品
全部の認定試験
F.A.Q
Databricks Machine Learning Professional 試験の主な試験内容は何ですか?-ktest
Databricks Machine Learning Professional 試験 になることに興味がある場合は、ktestが提供する最新のML Data Scientist認定資格 Databricks Certified Machine Learning Professional 試験ダンプを選択することを強くお勧めします。 これらの試験ダンプは、試験に簡単に合格できるように特別に設計されています。 すべての試験目的を包括的にカバーしており、試験の準備が万全であることが保証されます。 これらの Databricks Certified Machine Learning Professional 試験ダンプを使用すると、成功の可能性が高まり、自信を持って認定取得までの道のりを完了することができます。
Databricks Machine Learning Professional 試験の主な試験内容は何ですか?-ktest
Databricks Certified Machine Learning Professional 試験ダンプ
Databricks 認定機械学習エキスパート
Databricks Certified Machine Learning Professional 認定試験では、Databricks Machine Learning を使用する個人の能力と、運用タスクで高度な機械学習を実行する能力を評価します。 これには、機械学習実験を追跡、バージョン管理、管理する機能と、機械学習モデルのライフサイクルを管理する機能が含まれます。 さらに、認定試験では、機械学習モデルを導入するための戦略を実装する能力も評価されます。 最後に、候補者は、データのドリフトを検出する監視ソリューションを構築する能力について評価されます。 この認定試験に合格した個人は、Databricks Machine Learning を使用して高度な機械学習エンジニアリング タスクを実行することが期待されます。

Databricks Machine Learning Professional試験の詳細
タイプ: 査察証明書
質問数: 60 問の多肢選択問題
制限時間:120分
登録料: $200
言語: 英語
配信方法:オンライン傍聴
前提条件: なし。ただし、関連するトレーニングを強く推奨します。
推奨される経験: 試験ガイドラインに概要が記載されている機械学習タスクを実行する 1 年以上の実務経験
有効期間:2年
再認定: 認定ステータスを維持するには再認定が必要です。 Databricks 認定は、発行日から 2 年間有効です。

試験のテーマ
パート 1: 実験 - 30%
データ管理
● デルタテーブルの読み取りと書き込み
● デルタ テーブルの履歴を表示し、以前のバージョンのデルタ テーブルをロードする
● 機械学習ワークフローでの特徴ストレージ テーブルの作成、上書き、マージ、読み取り
実験の追跡
● MLflow を使用してパラメータ、モデル、評価指標を手動で記録します
● MLflow 実験のデータ、メタデータ、モデルにプログラムでアクセスして使用します。
高度な実験追跡
● モデル署名と入力例を使用して、MLflow 実験追跡ワークフローを実行します。
● ネストされた実行をトレースするための要件を決定する
● Hyperopt の使用を含む、自動ログを有??効にするプロセスについて説明します。
● SHAP ダイアグラム、カスタム ビジュアライゼーション、特徴データ、画像、メタデータなどのアーティファクトを記録および表示する

パート 2: モデルのライフサイクル管理 - 30%
前処理ロジック
● MLflow スタイルと MLflow スタイルを使用する利点について説明する
● pyfunc MLflow スタイルを使用する利点について説明する
● カスタム モデル クラスとオブジェクトに前処理ロジックとコンテキストを含めるプロセスと利点について説明する
モデル管理
● モデル レジストリの基本的な目的とユーザー インタラクションについて説明します。
● 新しいモデルまたは新しいモデルのバージョンをプログラムで登録します。
● 登録モデルおよび登録モデルのバージョンにメタデータを追加する
● 利用可能なモデル ステージを特定、比較、対比する
● モデル バージョンの変換、アーカイブ、削除
モデルのライフサイクルの自動化
● ML CI/CD パイプラインにおける自動テストの役割を決定する
● モデル レジストリ Webhook と Databricks ジョブを使用してモデルのライフサイクルを自動化する方法を説明する
● 汎用クラスタと比べてジョブ クラスタを使用する利点を判断する
● 特定のシナリオでモデルがステージ間を移行するときにトリガーされるジョブを作成する方法を説明する
● Webhook をジョブに接続する方法を説明する
● 表示された Webhook をトリガーするコード ブロックを決定する
● HTTP Webhook の使用例と Webhook URL の取得元を決定します。
● すべての Webhook を一覧表示する方法と Webhook を削除する方法について説明します。

パート 3: モデルの導入 - 25%
バッチ
● バッチ デプロイメントを、大部分のデプロイメント ユース ケースに対する適切なユース ケースとして説明する
● バッチ展開で予測を計算する方法を決定し、後で使用できるようにどこかに保存します。
● 事前計算されたバッチ予測をクエリすることによるリアルタイム サービスの利点を判断する
● 他のユースケースの解決策としてパフォーマンスの低いデータ ストアを特定する
● 登録されたモデルをロードするには、load_model を使用します。
● 単一ノード モデルを並列デプロイするには、spark_udf を使用します。
● Z ソートは、テーブルから予測を読み取る時間を短縮するためのソリューションと考えてください。
● 共通列のパーティションを特定してクエリを高速化します
●score_batch オペレーションを使用する実際的な利点について説明する
ストリームメディア
● ETL パイプラインの一般的な処理ツールとしての構造化ストリーミングについて説明する
● 構造化ストリーミングを受信データの連続推論ソリューションとして特定する
● 複雑なビジネス ロジックをストリーミング デプロイメントで処理する必要がある理由を説明する
● 構造化ストリーミングを通じて、データが順序どおりに到着しない可能性があることを判断します。
● 時間ベースの予測ストレージ内の継続的な予測をストリーミング展開シナリオとして認識します。
● バッチ デプロイメント パイプライン推論をストリーミング デプロイメント パイプラインに変換する
● バッチ デプロイメント パイプラインの書き込みをストリーミング デプロイメント パイプラインに変換する
リアルタイム
● 少数のレコードに対して、または高速な予測計算が必要な場合にリアルタイム推論を使用する利点について説明する
● リアルタイム展開の必要性としての JIT 機能の価値を判断する
モデルサービスのデプロイメントと各段階のエンドポイントについて説明する
● Model Service がモデルのデプロイメントに共通クラスターを使用する方法を決定する
● 本番およびステージング段階のクエリ モデル サービス対応モデル
● コンテナ内のクラウドによって提供される RESTful サービスが、本番グレードのリアルタイム デプロイメントにとってどのように最適なソリューションであるかを判断する

パート 4: ソリューションとデータ監視 - 15%
ドリフトタイプ
● ラベル ドリフトと特徴ドリフトを比較対照する
● 機能ドリフトやラベル ドリフトが発生する可能性があるシナリオを特定する
● コンセプトのドリフトとそれがモデルの有効性に及ぼす影響について説明する
ドリフトテストとモニタリング
● デジタル特徴ドリフトに対する簡単なソリューションとして概要統計モニタリングを説明する
● カテゴリカル特徴量ドリフトに対する簡単な解決策としてパターン、固有値、欠損値を説明する
● テストを、単純な要約統計よりも強力な数値的特徴ドリフト監視ソリューションとして説明する
● カテゴリ特徴量ドリフトを監視するための、単純な要約統計よりも強力なソリューションとしてテストを説明する
● 数値ドリフト検出のためのジェンソン?シャノン発散検定とコルモゴロフ?スミルノフ検定を比較対照する
● カイ二乗検定が役立つ状況を特定する
包括的なドリフトソリューション
● コンセプト ドリフトと機能ドリフトを測定するための一般的なワークフローを説明する
● いつ再トレーニングして更新されたモデルをデプロイするかを決定することが、ドリフト問題の解決策となる可能性があります
● 更新されたモデルのパフォーマンスが最新のデータで向上するかどうかをテストします。

Databricks Machine Learning Professional の無料問題集を共有します

1. Which of the following Databricks-managed MLflow capabilities is a centralized model store?
A.Models
B.Model Registry
C.Model Serving
D.Feature Store
E.Experiments
Answer: C

2. A machine learning engineer wants to log and deploy a model as an MLflow pyfunc model. They have custom preprocessing that needs to be completed on feature variables prior to fitting the model or computing predictions using that model. They decide to wrap this preprocessing in a custom model class ModelWithPreprocess, where the preprocessing is performed when calling fit and when calling predict. They then log the fitted model of the ModelWithPreprocess class as a pyfunc model.
Which of the following is a benefit of this approach when loading the logged pyfunc model for downstream deployment?
A.The pvfunc model can be used to deploy models in a parallelizable fashion
B.The same preprocessing logic will automatically be applied when calling fit
C.The same preprocessing logic will automatically be applied when calling predict
D.This approach has no impact when loading the logged Pvfunc model for downstream deployment
E.There is no longer a need for pipeline-like machine learning objects
Answer: E

3. Which of the following MLflow Model Registry use cases requires the use of an HTTP Webhook?
A.Starting a testing job when a new model is registered
B.Updating data in a source table for a Databricks SQL dashboard when a model version transitions to the Production stage
C.Sending an email alert when an automated testing Job fails
D.None of these use cases require the use of an HTTP Webhook
E.Sending a message to a Slack channel when a model version transitions stages
Answer: B

4. Which of the following lists all of the model stages are available in the MLflow Model Registry?
A.Development. Staging. Production
B.None. Staging. Production
C.Staging. Production. Archived
D.None. Staging. Production. Archived
E.Development. Staging. Production. Archived
Answer: A

5. A machine learning engineer needs to deliver predictions of a machine learning model in real-time. However, the feature values needed for computing the predictions are available one week before the query time.
Which of the following is a benefit of using a batch serving deployment in this scenario rather than a real-time serving deployment where predictions are computed at query time?
A.Batch serving has built-in capabilities in Databricks Machine Learning
B.There is no advantage to using batch serving deployments over real-time serving deployments
C.Computing predictions in real-time provides more up-to-date results
D.Testing is not possible in real-time serving deployments
E.Querying stored predictions can be faster than computing predictions in real-time
Answer: A

6. Which of the following describes the purpose of the context parameter in the predict method of Python models for MLflow?
A.The context parameter allows the user to specify which version of the registered MLflow Model should be used based on the given application's current scenario
B.The context parameter allows the user to document the performance of a model after it has been deployed
C.The context parameter allows the user to include relevant details of the business case to allow downstream users to understand the purpose of the model
D.The context parameter allows the user to provide the model with completely custom if-else logic for the given application's current scenario
E.The context parameter allows the user to provide the model access to objects like preprocessing models or custom configuration files
Answer: A

7. A machine learning engineering team has written predictions computed in a batch job to a Delta table for querying. However, the team has noticed that the querying is running slowly. The team has already tuned the size of the data files. Upon investigating, the team has concluded that the rows meeting the query condition are sparsely located throughout each of the data files.
Based on the scenario, which of the following optimization techniques could speed up the query by colocating similar records while considering values in multiple columns?
A.Z-Ordering
B.Bin-packing
C.Write as a Parquet file
D.Data skipping
E.Tuning the file size
Answer: E
emailsales@Ktest.jp
support@Ktest.jp
オフィス営業時間: 9:00 – 19:00 (月曜日から金曜日まで)
mixi  tumblr  facebook