RelFM
RelFM( *, database: Optional[str] = None, schema: Optional[str] = None, graph: Optional[Graph] = None, property_transformer: Optional[PropertyTransformer] = None, context: Optional[b.Relationship | b.Fragment | b.Chain] = None, validation: Optional[b.Relationship | b.Fragment | b.Chain] = None, source_concept: Optional[b.Concept] = None, task_type: Optional[str] = None, eval_metric: Optional[str] = None, use_current_time: bool = True, has_time_column: Optional[bool] = None, test_batch_size: Optional[int] = None, stream_logs: bool = True, dataset_alias: Optional[str] = None, parallel_reasoners_init: bool = True, n_estimators: int = 8, random_state: Optional[int] = 42, dfs_max_depth: int = 2, sample_size: Optional[int] = None, sampling_strategy: str = "stratified", device: Literal["cpu", "cuda"] = "cpu", clamp_min: Optional[int] = 0, clamp_max: Optional[int] = 100, regression_output: Literal["point", "distribution"] = "point", quantile_levels: Optional[List[float]] = None, norm_methods: Optional[Union[str, List[str]]] = None)Predict in-context with RelFM — a foundation model that requires no training.
Unlike GNN,
RelFM has a single workflow: construct it with context (used directly
as labeled in-context examples) and call RelFM.predictions — there is no
fit(), load(), or register_model().
Parameters
(databasestr, default:None) - Snowflake database to save predictions in.
(schemastr, default:None) - Snowflake schema to save predictions in.
(graphGraph, default:None) - The knowledge graph with edges defined. Used to pull in features from related tables via multi-hop DFS (seedfs_max_depth).
(property_transformerPropertyTransformer, default:None) - Column-level semantic type annotations. If omitted, all column types are auto-inferred.
(contextRelationship or Fragment, default:None) - Labeled split used directly as in-context examples (RelFM has no training step, so this plays the role of GNN’strainsplit at inference time).
(validationRelationship or Fragment, default:None) - Optional validation split.
(source_conceptConcept, default:None) - Source concept (inferred fromcontextif omitted).
(task_typestr, default:None) - One of"binary_classification","multiclass_classification", or"regression". Link prediction and multilabel classification are not supported by RelFM.
(eval_metricstr, default:None) - Evaluation metric compatible with the chosentask_type.
(use_current_timebool, default:True) - Use the current timestamp as the prediction time. Default isTrue.
(has_time_columnbool, default:None) - Set toTruewhen the task relationships use theatkeyword for temporal ordering.
(test_batch_sizeint, default:None) - Batch size used during inference. Default is256.
(stream_logsbool, default:True) - Stream logs to stdout. Default isTrue.
(dataset_aliasstr, default:None) - User chosen alias for the dataset.
(parallel_reasoners_initbool, default:True) - Initialize the Predictive and Logic reasoners in parallel. Default isTrue.
(n_estimatorsint, default:8) - Ensemble size. More = better quality, slower inference. Default is8.
(random_stateint, default:42) - Seed for ensemble generation and context sampling. Default is42.
(dfs_max_depthint, default:2) - How many hops across related tables to pull features from before running RelFM (1-3). Default is2.
(sample_sizeint, default:None) - Max context rows sampled fromcontextbefore feature generation. Default (None) uses all rows.
(sampling_strategystr, default:“stratified”) - How to draw the context sample:"stratified"(default),"most_recent","mixed","random", or"balanced".
(devicestr, default:“cpu”) - Inference device,"cpu"(default) or"cuda".
(clamp_minint, default:0) - Min percentile clamp for regression predictions. Default is0.
(clamp_maxint, default:100) - Max percentile clamp for regression predictions. Default is100.
(regression_outputstr, default:“point”) -"point"(default) or"distribution".
(quantile_levelslist of float, default:None) - Quantile levels output whenregression_output="distribution".
(norm_methodsstr or list of str, default:None) - Feature normalization method(s).
Examples
Assuming the setup from the module-level Quick Start
(relationalai.semantics.reasoners.predictive):
relfm = RelFM( graph=gnn_graph, property_transformer=property_transformer, source_concept=Students, context=Train, validation=Validation, task_type="binary_classification", eval_metric="roc_auc",)Students.predictions = relfm.predictions(domain=Test)Methods
.predictions()
RelFM.predictions(domain: b.Relationship | b.Fragment | b.Chain) -> b.RelationshipGenerate predictions on a test domain.
Materializes the context/validation splits together with the test
table, submits a prediction job, and returns a
Relationship that can be assigned to a
concept field for downstream querying.
The prediction attributes available on the returned relationship depend on the task type:
- Classification:
.probs,.predicted_labels— same shape as GNN’s classification output. - Regression:
.predicted_value(or, whenregression_output="distribution", one column per configuredquantile_levels).
Parameters:
(domainRelationship or Fragment or Chain) - The test split relationship (e.g. theTestrelationship defined during data modeling).
Returns:
Relationship- A prediction relationship to be assigned to the source concept (e.g.User.predictions = relfm.predictions(domain=Test)).
Raises:
TypeError- Ifdomainis not a Relationship, Fragment, or Chain.ValueError- If the test domain schema does not match thecontextschema.NotImplementedError- If deployments are enabled in the model configuration. Predictive reasoning is not yet supported in deployments.
.display_dataset_diagram()
RelFM.display_dataset_diagram(show_dtypes: bool = False) -> NoneDisplay the dataset schema inline in a Jupyter notebook.
Renders a diagram of the tables, columns, and foreign-key
relationships of the prepared dataset as an SVG in the current
notebook cell. Requires the Graphviz dot command-line tool
to be installed (brew install graphviz or
apt install graphviz).
Parameters:
(show_dtypesbool, default:False) - Include column data types in the diagram. Default isFalse.
Raises:
ValueError- If no dataset has been prepared yet (i.e.RelFM.predictionsorRelFM.scorehas not been called).
Examples:
relfm.predictions(domain=Test)relfm.display_dataset_diagram(show_dtypes=True).save_dataset_diagram()
RelFM.save_dataset_diagram(path: str, show_dtypes: bool = False) -> NoneSave the dataset schema diagram to a file.
Writes a diagram of the tables, columns, and foreign-key
relationships of the prepared dataset to path. The output
format is inferred from the file extension (e.g. svg,
png, pdf). Requires the Graphviz dot command-line
tool to be installed (brew install graphviz or
apt install graphviz).
Parameters:
(pathstr) - Destination file path, e.g."schema.svg"or"schema.png".
(show_dtypesbool, default:False) - Include column data types in the diagram. Default isFalse.
Raises:
ValueError- If no dataset has been prepared yet (i.e.RelFM.predictionsorRelFM.scorehas not been called).
Examples:
relfm.predictions(domain=Test)relfm.save_dataset_diagram("schema.svg").score()
RelFM.score(domain: Optional[b.Relationship | b.Fragment | b.Chain] = None) -> dictScore RelFM against a labeled domain and return evaluation metrics.
Unlike RelFM.predictions, score() needs true labels to compare
against — pass a labeled split with the same shape as context
(source concept plus label column), or omit domain to score
against the validation split already configured on this instance.
Runs the same DFS + ICL pipeline as RelFM.predictions, but against
this split, then computes every metric available for the task type —
not just the configured eval_metric — since RelFM has no training
loop to justify tracking only one.
Parameters:
(domainRelationship, Fragment, or Chain, default:None) - A labeled split to score against. Defaults to thevalidationsplit passed to the constructor.
Returns:
dict-{"metric_name": str, "metrics": dict[str, float]}— the configuredeval_metric’s name (also a key inmetrics), plus every metric available for this task type.
Raises:
ValueError- If no domain is available (neither passed nor configured viavalidation=at construction), or ifdomainis not a Relationship, Fragment, or Chain.NotImplementedError- If deployments are enabled in the model configuration.