August 15, 2023

Scenario: Indexing a reference data set in Elasticsearch – Docs for ESB 6.x

Scenario: Indexing a reference data set in Elasticsearch

This scenario applies only to a subscription-based Talend Platform solution with Big data or Talend Data Fabric.

In this Job, the tMatchIndex component creates an index in
Elasticsearch and populates it with a clean and deduplicated data set which contains a
list of education centers in Chicago.

After performing all the matching actions on the data set which contains a list of
education centers in Chicago, you do not need to restart the matching process from
scratch when you get new data records having the same schema. You can index the clean
data set in Elasticsearch using tMatchIndex for continuous
matching purposes.

Before indexing a reference data set in Elasticsearch:

  • You generated a pairing model using tMatchPairing.

    You can find examples of how to generate a pairing
    model on Talend Help Center (https://help.talend.com).

  • Make sure the input data you want to index is clean and deduplicated.

    You can find an example of how to clean and
    deduplicate a data set on Talend Help Center (https://help.talend.com).

  • The Elasticsearch cluster must be running Elasticsearch 5+.


Document get from Talend https://help.talend.com
Thank you for watching.
Subscribe
Notify of
guest
0 Comments
Inline Feedbacks
View all comments
0
Would love your thoughts, please comment.x
()
x