NGワード絞り込みスレッド★119

[過去ﾛｸﾞ] NGワード絞り込みスレッド★119 (655ﾚｽ)
上下前次1-新
抽出解除必死ﾁｪｯｶｰ(本家) (べ) 自ID ﾚｽ栞あぼーん

このｽﾚｯﾄﾞは過去ﾛｸﾞ倉庫に格納されています｡
次ｽﾚ検索歴削→次ｽﾚ栞削→次ｽﾚ過去ﾛｸﾞﾒﾆｭｰ

599: YAMAGUTIseisei (ﾜｯﾁｮｲ 5e48-baM0) 2018/08/27(月)10:00 ID:sC3a+vHE0(1/2) BE AAS
BEｱｲｺﾝ:nida.gif
Other methods of exploration are designed to work in combination with maximizing a reward function, such as those utilizing uncertainty about value function estimates [5, 23], or those using perturbations of the policy for exploration [8, 29].
他の探査方法は、価値関数推定値に関する不確実性を利用する報酬関数や探索のための方針の摂動を用いる報酬関数などの報酬関数を最大化することと組み合わせて機能するように設計されている[8]、[29]。
Schmidhuber [37]とOudeyer [25]、OudeyerとKaplan [26]は、内在的動機づけへのアプローチに関する初期の研究のいくつかについて素晴らしいレビューを提供する。
Alternative methods of exploration include Sukhbaatar et al.
探査の代替方法には、Sukhbaatar et al。
[45] where they utilize an adversarial game between two agents for exploration.
省6

600: YAMAGUTIseisei (ﾜｯﾁｮｲ 5e48-baM0) 2018/08/27(月)10:00 ID:sC3a+vHE0(2/2) BE AAS
BEｱｲｺﾝ:nida.gif
In Gregor et al.
Gregor et al。
[10], they optimize a quantity called empowerment which is a measurement of the control an agent has over the state.
[10]、エージェントはエンパワーメントと呼ばれる量を最適化します。これは、エージェントがその状態を超えた制御の測定値です。
In a concurrent work, diversity is used as a measure to learn skills without reward functions Eysenbach et al.
並行作業では、報酬機能なしにスキルを習得するための手段として多様性が使用されます。Eysenbach et al。
省7

上下前次1-新書関写板覧索設栞歴

ｽﾚ情報赤ﾚｽ抽出画像ﾚｽ抽出歴の未読ｽﾚ AAｻﾑﾈｲﾙ

ぬこの手ぬこTOP 0.026s