window parameter (see
below) with a value of “all”.
We use gensim to train the item2vec model, so for further details also see it’s
word2vec page.
Usage
The following example shows how the step can be used in a recipe.Examples
Examples
- Example 1
- Signature
The following uses default parameter values only, and thus would be equivalent to using the step without specifying
any parameters.
Inputs & Outputs
The following are the inputs expected by the step and the outputs it produces. These are generally columns (ds.first_name), datasets (ds or ds[["first_name", "last_name"]]) or models (referenced
by name e.g. "churn-clf").
Inputs
Inputs
Outputs
Outputs
column[list[number]]
required
A list column containing item embeddings in the same order as the items input column. Embeddings are lists of numbers
(vectors).
Configuration
The following parameters can be used to configure the behaviour of the step by including them in a json object as the last “input” to the step, i.e.step(..., {"param": "value", ...}) -> (output).
Parameters
Parameters
integer
default:"48"
Length of resulting embedding vectors.Values must be in the following range:
integer
default:"1"
Whether to use the skip-gram or CBOW algorithm.
Set this to 1 for skip-gram, and 0 for CBOW.Values must be in the following range:
integer
default:"20"
Update maximum for negative-sampling.
Only update these many word vectors.
number
default:"0.025"
Initial learning rate.Values must be in the following range:
[integer, string]
default:"5"
Size of word context window.
Must be either an integer (the number of neighbouring words to consider), or any of “auto”, “max” or “all”,
in which case the window is equal to the whole list/session/basket.
Options
Options
- integer
- string
integer
integer.Values must be in the following range:
integer
default:"3"
Minimum count of item in dataset.
If an item occurs fewer than this many times it will be ignored.Values must be in the following range:
integer
default:"10"
Iterations.
How many epochs to run the algorithm for.Values must be in the following range:
number
default:"0"
Percentage of most-common items to filter out (equivalent to “stop words”).Values must be in the following range:
boolean
default:"true"
Whether to return normalized item vectors.