TDID Sample Questions & Answers
Four areas share the top weight: handling files with tMap, joining and filtering records, connecting to databases, and sequencing jobs end to end, next to integration basics, context variables, error handling, deployment, project organization and debugging.
Launch the full TDID simulator →Showing 10 of 20 free samples.
- Question 1Advanced
Working with Databases · Transaction Management with Parent/Child Jobs
Case Study:
A retail company, StyleSphere, needs to build a daily job to process sales data. The job must first download a ZIP file from an FTP server. This ZIP file contains three CSV files:
products.csv,sales.csv, andstores.csv. After unzipping, the job must load the data from all three files into corresponding tables in a PostgreSQL database. The entire process must be transactional; if loading any of the three files fails, all changes made to the database during that run must be rolled back.The development team has decided to use a parent job for orchestration. This parent job will handle the FTP download, unzipping, and database connection management. It will then call a child job to perform the actual loading of the three files.
Requirements:
- The database connection must be opened once in the parent job and shared with the child job.
- The final database commit should only happen in the parent job after the child job completes successfully.
- If the child job fails, the parent job must execute a database rollback.
Which design pattern correctly fulfills all these requirements?
Show answer & explanation
Correct answer: C
This is the correct and standard Talend pattern for managing transactions across parent and child jobs. 1) The
tDBConnectionin the parent with auto-commit off starts the transaction. 2) The child job uses this shared connection via the 'Use or register a shared DB Connection' option intRunJob. 3) ThetDBOutputcomponents in the child job write data within this single transaction. 4) The parent job regains control and, based on the success or failure of thetRunJobcomponent, executes either atDBCommitor atDBRollback, thus managing the entire unit of work transactionally. - Question 2Beginner
Working with Files · Data Flow Merging
A developer uses a
tUnitecomponent to merge data from three different source files that have identical schemas. However, during execution, the job fails with a schema mismatch error. What is the most likely cause of this error?Show answer & explanation
Correct answer: B
The
tUnitecomponent strictly requires that all incoming data flows have the exact same schema structure, including column names (which are case-sensitive), data types, precision, and length. Even a minor difference, like 'CustomerID' vs 'customerid', or aString(10)vs aString(12), will be considered a schema mismatch and cause the component to fail. The developer must ensure all input schemas are identical, often by using a single Repository schema for all inputs. - Question 3Beginner
Getting Started with Data Integration · Metadata Management
When building a Talend job, a developer needs to define the structure of a data source once and reuse it across multiple components and jobs. Which Talend feature is designed for this purpose?
Show answer & explanation
Correct answer: C
Repository Metadata is the central feature in Talend for storing reusable information about data sources, such as file schemas, database connections, and table definitions. By defining a schema in the Metadata section of the Repository, a developer can drag and drop it onto components, ensuring consistency and ease of maintenance. If the source structure changes, updating the central Repository schema can propagate the change to all jobs that use it.
- Question 4Intermediate
Joining and Filtering Data · Conditional Data Routing
A Talend job processes customer data and needs to perform different actions based on the customer's country. The logic should be: if the country is 'USA', load to a specific target; if the country is 'CAN', load to another target; for all other countries, log the record and discard. What is the most efficient way to implement this routing logic in a single component?
Show answer & explanation
Correct answer: B
The
tMapcomponent is ideal for complex routing logic. It allows for the creation of multiple named output flows from a single input. Each output can have a filter expression applied directly to it. This allows the developer to define conditions (e.g.,row1.country.equals("USA")) for each path.tMapalso supports a 'catch output reject' option for rows that don't match any filter, which is perfect for the logging requirement. This is more efficient and cleaner than using multipletFilterRowcomponents. - Question 5Intermediate
Orchestrating Jobs · Parent-Child Job Communication
A developer needs to pass a value, specifically a record count calculated in a child job, back to its parent job for logging purposes. Which combination of components and settings is the standard method to achieve this?
Show answer & explanation
Correct answer: C
This is the designed pattern in Talend for returning a data flow (even a single row with a single value) from a child job to a parent. The
tBufferOutputin the child job writes the data to an in-memory buffer. The parent job'stRunJobcomponent can be configured with a schema that matches thetBufferOutput, which then provides a standardMaindata flow output link that can be connected to subsequent components in the parent job. - Question 6Beginner
Orchestrating Jobs · Routines vs Joblets
What is the primary difference between a Routine and a Joblet in Talend Studio?
Show answer & explanation
Correct answer: B
This is the core distinction. A Routine is a collection of static Java methods that can be called from any expression editor in Talend (like in a
tMap). They are for code-level reusability. A Joblet, on the other hand, is a reusable piece of a job's graphical design. It encapsulates a sequence of components and their connections, allowing for the reuse of entire data processing steps across multiple jobs. - Question 7Intermediate
Orchestrating Jobs · Parallelization
A project requires processing a 10 GB file containing millions of records. The processing for each record is independent but computationally intensive. To speed up the job, the lead developer suggests using parallelization. Which Talend component is specifically designed to split a data flow into multiple parallel threads for processing?
Show answer & explanation
Correct answer: D
The
tParallelizecomponent is the standard mechanism in Talend Data Integration for achieving data flow parallelization. It takes an input flow and creates a specified number of concurrent sub-processes (threads) to execute a defined part of the job. It then synchronizes these threads before passing the combined results to the next component. This is ideal for improving performance on multi-core machines when processing large volumes of data where individual records can be handled independently. - Question 8AdvancedSelect 2
Error Handling · Centralized Error Handling Framework
A developer needs to implement a global error handling mechanism for a project. The goal is to capture any Java exception or component failure (
tDie) from any job, log the error details (job name, component, error message) to a database table, and send an email notification. Which combination of components provides a centralized, reusable solution for this? (Select TWO)Show answer & explanation
Correct answers: A, C
- Question 9Beginner
Joining and Filtering Data · tJoin Component Behavior
True or False: When using a
tJoincomponent, if a row from the main input flow does not have a matching row in the lookup flow, that row will be discarded and will not appear in the output, regardless of the join type selected.Show answer & explanation
Correct answer: B
This statement is false. The behavior described is that of an 'Inner Join'. However, the
tJoincomponent also supports 'Left Outer Join'. If Left Outer Join is selected, all rows from the main input flow will be included in the output. If a match is found in the lookup, the lookup columns will be populated; if no match is found, the lookup columns will be populated with null values. - Question 10Advanced
Orchestrating Jobs · Job Design for Performance and API Limits
Case Study:
An e-commerce company, GearUp, is building an end-of-day reporting job. The job must process a large CSV file of transactions. For each transaction, it needs to look up product details from a PostgreSQL database and customer details from a Salesforce instance. The final output should be an aggregated summary report written to an Excel file.
Current Situation:
- The transaction CSV file has 5 million rows.
- The product table in PostgreSQL has 100,000 rows.
- The customer data in Salesforce is accessed via API and has a rate limit of 10,000 calls per hour.
- The initial job design uses a single
tMapcomponent to perform both the PostgreSQL lookup and the Salesforce lookup for each transaction row.
Problem:
The job is extremely slow and frequently fails due to Salesforce API rate limiting. The company needs a more robust and performant design.Which redesigned approach would best address the performance and rate-limiting issues?
Show answer & explanation
Correct answer: C
This design pattern is the most effective solution. It addresses the core problem by dramatically reducing the number of API calls. Instead of 5 million individual API calls, it makes a few bulk calls to fetch only the necessary customer data. Caching this data locally (in a database table or in-memory using
tHashcomponents) transforms the slow, rate-limited API lookup into a fast, local lookup. This decoupling and pre-caching strategy is a best practice for dealing with API-based lookups in high-volume ETL jobs.
Ready for the real thing?
The full TDID simulator has every exam-style question, timed mode, and instant scoring.