DataStitch :: Product Help

The Approach

Datasets and Records
​A dataset is a set of records. All functions in DataStitch are detaset-centric. Datasets are automatically created from an import of CSV files along with the information about it. In future versions, datasets will be created by operations on existing ones. In DataStitch's datasets, the source data is maintained separately from its indexes which are required to determine duplicates. 
CSV Files, Data Fields and Headers
​CSV files are the inputs to DataStitch. An import of a CSV file results in a dataset. The CSV can contain any number of fields. It is of paramount importance that the imported CSV file has a header. Without one, DataStitch will not be able to perform any field level processing on the resulting dataset. The DeDuplication feature is dependent on field names (see section: DeDuplication Rules).

If for some reason the import process is aborted or unable to complete, the application will not creates a dataset. 
Indexing Rules
​The CSV field headers are very important in DataStitch as they determine how the fields are indexed. It is recommended that right field names are given to the fields before importing into DataStitch. 

Future versions of DataStitch will provide a functionality to edit the field headers. For now, if you have imported a file and the field name was supposed to be different, you will need to delete the dataset and change the header before importing again. 

The rules of field names are as follows:
  • The 1st three fields with the word “Name” will all be considered as Person's Name. Examples are: First Name, MiddleName, Father-Name, Mother-Maiden-Name, NameOfSpouse, ChildName1, etc. The presence of string "name" in the field header will make this assignment. 
  • The 1st field with the word “Company” will be considered as Company or Organization's Name. We need to be careful here to not to have the word "name" in this field, else it will be treated as a person's name. A good practice is to keep the field header of the name of organization or company as "Company".  
  • The 1st three fields with the word “Date” will all be considered as Date fields. Examples are: Registration Date, WarrantyDate, Birthdate, DateOfExpiry, etc. 
  • The 1st five fields with the word “Text” will all be considered as compressed strings. Add the string "text" to the header of the field if you want it to be indexed as a text field. For example, if the field IC_Number or ReferenceCode need to be indexed, they need to be renamed with the word "text" in their names. Resulting names could be IC_Number_Text or TextReferenceCode. The text fields are indexed on all characters except spaces. Hence, a field called PassportNoText with record values of "H 144234", "h 14 42 34" and " H 1 4 4 2 3 4 " are considered as same during indexing. 
  • The 1st field with word “Mobile” will be treated as Mobile Number. 
  • The 1st field with word “Phone” will be treated as Phone Number.
  • The 1st field with word “Email” will be treated as Email Address.
  • The 1st field with the word “URL” will be considered as URL.
  • The 1st three fields with the word “Address” will all be considered as Address. These are usually the Street and location address fields. Examples are: StreetAddress-1, Address-2, etc
  • The field names of City, State, Country and ZIP will be treated as their meaning in address fields. If you are using PostalCode as the field name and you wish it to be indexed, please change the field name to PostalCodeZIP. 
Additional Notes on Indexing
  • A field can be indexed in only one way by using DeDuplication Rules explained above.  
  • The field of "State" cannot be indexed without the presence of a "Country" field.
  • If Firstname and Lastname are 2 separate fields, then they will be indexed separately. If Firstname+Lastname need to be deduplicated then they both need to be selected. Alternatively, one can load the field as "Fullname" with the Firstname and Lastname combined. 
  • The text field indexes only the alphanumeric characters in it. Hence the following 3 field values in a text field: "A1234", " a 12 3 4 " and "(A)12-34*" will all be treated as same duplicate fields. 
Data Security
​DataStitch runs on a user's computer. There is no data that is sent from DataStitch to any 3rd party server or vice-versa. Hence, the application can be run on computers which are not connected to a network. 

The CSV files which are used to create datasets are never changed. The output files are maintained in a separate folder with a different filename. 

Product Usage

Installation Steps
DataStitch is condensed in a single executable file. It is downloaded as a compressed file (DataStitch.zip) and runs from a folder. The database is maintained by the application and no intervention is needed. 

1. Locate a space on your computer where DataStitch will be installed. Depending on your data, please ensure that there is sufficient storage at this location. While DataStitch does not modify your original data files, it makes a copy of it within its database for indexing and processing. 

2. Download the latest version of DataStitch from the link sent to you in the email. Place it in the located folder and unzip it in that folder. DataStitch.zip will create a folder called DataStitch.

3. Go to DataStitch folder and double-click INSTALL.bat. This will prepare the database environment. This is to be done only once. 

4. Run DataStitch.bat. After a few seconds, the application will start and is ready to use. Note that it will run a database server in a separate window. You will need to leave it running in the background. That window will close automatically when you exit DataStitch. 

5. It is recommended that you create a shortcut to DataStitch.bat file and add it to a convenient location. If the application is invoked from the shortcut then there is no need to come back to DataStitch folder again. The original file: DataStitch.zip can be archived. 
​
You are ready to use DataStitch. There are two setups that need to be done, which are explained in the product help page: 1) Configuration of folders for incoming and outgoing data and 2) Addition of license keys. 
Application Setup
There are two setup steps which are explained in this document itself. 
Refer to :
  1. Folder Configuration
  2. Managing License Keys 
Create Datasets by Importing CSV Files
A data file with fields in comma separated values can be imported into the application as a dataset.  
  • The CSV file MUST have headers. This is of importance. All subsequent operations on dataset are based on field headers. 
  • The field headers must comply with the indexing rule definitions. This will determine on how the data in that field will be indexed. 
  • The delimiter must be a comma (","). In future versions, this can be changed. 
  • The CSV file should be UTF-8 encoded.
  • The extension of the data file must be ".csv".

If the number of data columns in each row of CSV file do not match the number of header fields, DataStitch will give an import error and the dataset will not be created. 

There are two ways of importing a CSV file: The first is to enter the name and full path of the individual file which is on your system or network (eg, c:\campaigns\winter19.csv). The second and more common way is select them from a configured folder. All new CSV files can be added to this folder and they will start to appear in the list which can be selected in the application. 
Listing Datasets
​This function lists the datasets that have been successfully imported and indexed in the system. The original CSV file from which the dataset is created is never modified. By default, the dataset name is same as the CSV filename. The number of records shown does not include the first header record. 
Browsing Datasets
​The current version of DataStitch has basic functionality to browse the first 10, last 10 and random 10 records of a selected dataset. In future versions, this feature will have more advance ways of data browsing. 
Analyzing Datasets
When a dataset is created by DataStitch, it profiles it too. The option of 'Analyze' shares the selected dataset's profile. 
It has the following information:
  • Number of records in dataset, excluding the header
  • Field names of the dataset and whether they are indexed or not. If the field is indexed then other operations like clustering, merging and dataset operations can be performed on that field. For every field, the number of blank/Null values are also computed, displayed as an absolute number and a percentage within the dataset. 
  • Completeness of dataset, with the number of records containing non-blank values in each of the field.
Renaming and Deleting Datasets
​A dataset's name can be changed by 'Rename' option. 

The 'Delete' option removes the entire dataset from its database. If you have deleted a dataset by mistake then you will have to import it again from the original CSV file. 
Exporting Datasets
​The 'Export' option will take the content of the selected dataset and export it as a CSV file. The output is in the configured folder under Settings>Config. 

Sufficient balance of license keys is required to export a dataset. You will not be able to export a dataset if it's number of records are greater than the balance of license keys. 

Once the dataset is successfully exported, the balance key value is updated accordingly. 
Merging Similar Records
​Merging unifies multiple records into a single master record based on a common field value. The selected common field must be one which is already indexed. A single field can be selected. If merging is required based on a combination of fields, they need to be clustered prior to merging. 

The similar records are merged into a Master Record (MR). The non-empty fields of other records replace the empty fields in the MR. You can select the MR based on one of the following 4 criteria:
- The record with maximum amount of data (selected by default)
- The record with maximum number of non empty fields
- The first record in dataset
- The last record in dataset

Once MR is selected, the other records in that similar set will merge into Master. Note that the original dataset is not modified as the merged data is created in a new dataset. 
Clustering Similar Records
​Clustering creates common values for any combination of fields in any order. It creates a new dataset with one additional field of Cluster, which has the value of field combination and their order. The new clustered dataset can be exported along with the Cluster field for analysis, or used by Merge function to collapse multiple records into a single one based on identical Cluster field. Note that the empty or NULL values in a field will result in a similar Cluster field. 
Removing Exact Duplicates
This function searches for identical records in the dataset and removes them. The match is not on a particular field but the complete record. No new dataset is created.  

Say, a dataset of 100 records is found to have 2 duplicate clusters of 5 and 10 records respectively. This function will remove 4 records from the first cluster and 9 records from the second. The final dataset will have 87 records. 
Dataset Operations
Dataset Operations are powerful processing operations between two selected datasets.  Dataset Operations between two datasets take place on a single common indexed field. If the operation is needed on more than one indexed field then it needs to be clustered first. 

Three operations are possible:
  1. Union – All records of first and second dataset. 
  2. Intersection – Common records between first and second dataset, based on any order of indexed fields.
  3. Subtraction – Records in one dataset which are not in other, eg first dataset minus second. 
Folder Configuration
Under the Options section, two special folders can be configured. The first folder is to map the location on your computer or network which has the CSV files that need to be imported. This is used by the function of importing from configured folder. The second folder is to define where the exported datasets will be stored. 

It is important to specify the complete absolute path of the folders, ending with a "\".  For example: "D:\CRM\IncomingData\" or "c:\reports\exported_data\".
Managing License Keys
Keys will be needed to export the data from DataStitch. When 1000 records (say) are exported, 1000 keys are used. This option allows you to add keys to your application so that data can be exported. 

You will receive the keys from Contactous or your reseller in an email. It is a single encrypted string containing the recipient's email address and number of keys. The easy way is to copy this string from email and paste it in the DataStitch's option of 'License Key'. If the string is valid, it will reflect in the application as the number of keys. ​

​For evaluation version, 10,000 complementary keys can be requested. 
Zoho CRM Export Setup
The following steps will guide you to connect your DataStitch instance to Zoho CRM:

1. Select the Config option for access token generation and setting up configurations.
2. Select the 4th option for Zoho CRM. You will be redirected to a web browser for the Zoho authentication process. Enter your Zoho credentials here and click Login.
3. After logging in, you will be redirect to your CRM environment where the data will exported to. Here, select the right option and click the submit button.
4. Once you have selected a CRM, you will be redirected to a permissions section. Please check the permissions field and accept them to generate a Zoho refresh token.
5. After you accept the permission, you will be redirected to a page of DataStitch, where the refresh token can be seen. You just need to copy it into the clipboard.
6. Switch to the open DataStitch application and paste the copied refresh token there.
7. Datastitch will generate the access token automatically based on the refresh token. Now you are ready to export data to Zoho CRM from the Export option. 

For better illustration, the following video will help:

Product Information

DataStitch Roadmap
DataStitch's roadmap is driven by its users. We compile their inputs and feedback to add new features to the system. 

The following features are planned in upcoming versions: 
  • CSV Header Fields Mapping
  • Data Cleansing Routines
  • Optimizing memory and performance in dataset operations for >50,000 records.
Current Limitations
For large datasets (>10,000 records), the Merging and Removing Exact Duplicates is slow. This will be addressed in version 2.2, which is planned in November '21. 
Known Issues/Bugs
Non English characters are not correctly displayed in Dataset Browsing Mode.

Note - While DataStitch has been rigorously tested with large datasets and the occurrence of a crash is unlikely. But if it happens then a log file called php_errors.txt will be created in the root directory of DataStitch (where README file is located). We request you to send this file to support@contactous.com. 
Support
Sending a mail to support@contactous.com is the fastest way. Alternatively, use the web form and we will get back to you. ​

Contact Support