WORK WITH ZEKI

Frequently Asked Questions

We offer a variety of data delivery methods, including flat files and custom reports. Flat files can be delivered using Amazon S3, Snowflake, or via a link containing a zipped version of your flat file. Our most popular delivery method is through an Amazon S3 bucket where we can deliver parquet or CSV files to our clients.

You can access your Amazon S3 bucket using a pre-signed URL. This URL will be provided to you upon successful purchase and will grant you time-limited access to your data. The pre-signed URL is valid for a specified period (e.g., 24 hours). If the URL expires, you will need to request a new one. Each dataset will have a unique pre-signed URL to ensure secure and controlled access.

Clients will receive updated data on a quarterly basis. For example, data for Q1 (including headcounts, inflows, outflows, etc.) would become available by April 15th. If you would like your data to be updated more frequently, we also offer monthly delivery for select feeds.

Our data is gathered from a diverse range of publicly accessible online sources. These include academic publications, awards, and various other online platforms. This approach ensures that we provide comprehensive and up-to-date information.

Yes, if there are any specific companies you are interested in tracking that are not included in the standard datasets or sample files, we can arrange to develop a custom report upon request, this will be subject to the context of the companies and specifics you wish to investigate.

We do not associate subsidiary companies with their parent company. For instance, we would consider Alphabet as a standalone entity and not include Google or other subsidiaries like Google DeepMind under its umbrella. Google DeepMind would be treated as a separate company.

Yes, we can provide data feeds with custom granularities or mappings based on your requirements. This process will involve generating a detailed report to establish the new granularities.

We provide historical data for innovators between 2010-2025. Each data point along this time series is created from historical data i.e. publication information for that specified timeframe. This data is not revised retroactively or adjusted. This ensures that the data reflects the information available at that specific moment. Our historical data, including signal outputs like Zeki scores, is produced using only the information that would have been available at the time the signal is associated with. This maintains the integrity and accuracy of historical analyses. Our algorithms are trained on a diverse set of high-quality data sources to ensure robustness and accuracy. We take the risk of overfitting very seriously and employ various techniques, such as crossvalidation and regularization, to mitigate this risk and ensure our models generalize well to new data. Our algorithms are regularly reviewed and updated to incorporate new data inputs and human in the loop feedback. This continuous improvement process ensures that our scoring system remains accurate and relevant over time.

We do not report on data in China or include internal employee movements within companies. Not all work profiles are updated during job transitions. To clarify time series data, we weight metrics based on the total employees at a company during that period. We’ve developed a proprietary role taxonomy for Research and Development in Deep Tech Sectors and map job titles using custom machine learning algorithms. We create job title representations with a custom encoder to identify role clusters, including only those with 95% accuracy. We predict gender using first names, achieving 95% accuracy.

Can’t find an answer to your question?

Ask us anything, get in touch with Zeki.