DATA-ARCHITECT Sample Questions & Answers
Built around two tied leaders: persisting data in standard and custom objects, and scalable data model design, plus archiving and performance at large data volumes, migration techniques, governance and stewardship, and consolidating records company-wide.
Launch the full DATA-ARCHITECT simulator →Showing 10 of 20 free samples.
- Question 1Intermediate
Master Data Management · Data Survivorship
FinServe Inc. is implementing a Master Data Management (MDM) strategy for their customer data, which is sourced from their Salesforce org, a marketing automation platform, and a billing system. The goal is to create a 'golden record' for each customer in Salesforce. The billing system is considered the most reliable source for customer address information, while the marketing platform has the most up-to-date email addresses.
Which MDM concept should be configured to ensure the final golden record uses the address from the billing system and the email from the marketing platform?
Show answer & explanation
Correct answer: C
Data Survivorship Rules are the specific configurations that determine which data from which source system should be preserved in the final 'golden record' when duplicates are merged. In this case, a rule would be set to prioritize the 'address' field from the billing system and the 'email' field from the marketing platform.
- Question 2Intermediate
Data Modeling / Database Design · Data Skew
A data architect at Global Motors has identified significant account data skew on their
Service_Request__ccustom object, where a few global fleet accounts own millions of service request records. This is causing slow report performance and record locking issues during data updates. The relationship to the Account is a lookup.Which is the most appropriate initial step to mitigate the performance impact of this lookup skew?
Show answer & explanation
Correct answer: D
This is a common and effective strategy for mitigating lookup skew. By creating multiple 'bucket' or 'sub-accounts' under the main fleet account and distributing the child records among them, you reduce the number of children linked to any single parent record. This breaks up the skew and significantly reduces lock contention and improves performance.
- Question 3Intermediate
Salesforce Data Management · Data Quality and Duplicate Management
A company is experiencing low user adoption of their Salesforce CRM due to poor data quality. A data quality assessment reveals high rates of duplication, incompleteness, and inaccuracy in the Lead and Contact objects. The Data Architect has been tasked with recommending a solution to improve data quality at the point of entry.
Which declarative feature should be the primary recommendation to prevent the creation of new duplicate records?
Show answer & explanation
Correct answer: B
This is the standard, declarative feature set designed specifically for identifying and managing duplicates. A Matching Rule defines how to identify duplicates (e.g., based on fuzzy name and exact email), and a Duplicate Rule determines what to do when a duplicate is found (e.g., block creation or allow with an alert). This is the primary tool for preventing new duplicates.
- Question 4Advanced
Large Data Volume Considerations · Skinny Tables
Apex Innovations runs complex reports on their
OpportunityandOpportunityLineItemobjects, involving numerous formula fields and joins across 5 related custom objects. With over 20 million Opportunity records, these reports are frequently timing out. The most commonly used fields in the report filters and columns are spread across these objects.What Salesforce feature should the architect propose to specifically address this report performance issue?
Show answer & explanation
Correct answer: D
Skinny tables are the ideal solution for this scenario. They are custom database tables that contain a subset of fields from a standard or custom object, including fields from related objects. By combining fields from
Opportunity,OpportunityLineItem, and the related custom objects into a single, denormalized table, Salesforce can avoid expensive joins at runtime, dramatically improving report and query performance. - Question 5Advanced
Data Migration · Large Scale Migration Strategy
Quantum Corp, a large manufacturing firm, is migrating its legacy ERP data into a new Salesforce implementation. The migration involves approximately 100 million records across 15 related objects, including
Account,Product2,Asset, and several custom objects forWork_Order__candMaintenance_Plan__c. The business has mandated a maximum downtime window of 12 hours over a single weekend for the final data cutover.Constraints & Requirements:
- The legacy system must remain operational until the final cutover begins.
- Data integrity and all complex relationships must be perfectly preserved.
- The migration must be completed within the 12-hour window.
- Performance of the new org must not be degraded post-migration.
Which data migration strategy should the architect propose?
flowchart TD subgraph Pre-Cutover Phase A[Extract & Transform Legacy Data] --> B{Initial Bulk Load} B --> C[Suspend Automation & Sharing Rules] C --> D[Load Parent Objects e.g., Accounts] D --> E[Load Child Objects e.g., Assets] end subgraph Cutover Window (12 hours) F[Legacy System Offline] --> G{Delta Load} G --> H[Run Validation Scripts] H --> I[Re-enable Automation & Sharing Rules] I --> J[Perform Sharing Recalculation] end subgraph Post-Cutover K[New Salesforce Org Live] end E --> F J --> KShow answer & explanation
Correct answer: C
This is the correct and standard approach for large-scale migrations with tight downtime constraints. By pre-loading the majority of the data beforehand, the work performed during the critical 12-hour window is minimized to only the delta. Deferring sharing calculations is a key performance optimization that saves significant time during the load process, allowing the recalculation to happen after the data is in place.
- Question 6Intermediate
Data Modeling / Database Design · Person Accounts vs Custom Model
A wealth management firm is designing a data model for its high-net-worth clients and their households. The firm needs to track individual clients, the households they belong to, and the relationships between clients within a household (e.g., spouse, child). They also need to roll up financial summaries from individual clients to the household level.
Which data model provides the most standard, scalable, and feature-rich solution for this requirement?
Show answer & explanation
Correct answer: B
This is the correct, standard approach, especially in industries like financial services. Person Accounts are designed for B2C scenarios. The standard householding model (often part of Financial Services Cloud, but the principles apply) allows grouping Person Accounts into households. The Account Contact Relationship object is specifically designed to define the nature of relationships (e.g., 'Spouse', 'Dependent') between people and accounts, providing a complete, out-of-the-box solution.
- Question 7AdvancedSelect 2
Salesforce Data Management · Multi-Org Data Consolidation
A multinational corporation has separate Salesforce orgs for its North America and Europe business units. They need to provide their global executive team with a unified, 360-degree view of top-tier customer accounts, including opportunities and cases from both orgs, without forcing users to log into two different systems. The solution should avoid creating a single, massive global org at this time.
Which TWO solutions could an architect propose to meet this requirement? (Select TWO)
Show answer & explanation
Correct answers: B, C
This is a primary use case for Salesforce Connect. It allows you to create a central view of data from multiple orgs in real-time without physically moving or replicating the data. Executives can log into one org and see a unified, though virtualized, view.
This is a robust and common architectural pattern for multi-org consolidation. A middleware tool (like MuleSoft) orchestrates the flow of data, replicating key information into a central org. This provides better reporting capabilities and performance than pure virtualization, at the cost of data replication.
- Question 8Intermediate
Large Data Volume Considerations · PK Chunking
A developer is using the Bulk API to extract 15 million records from the Case object. The initial query is timing out due to the large data volume. The developer needs a reliable way to break the query into smaller, manageable chunks without writing complex logic to calculate ranges on non-indexed fields.
Which feature of the Bulk API should be used to solve this issue?
Show answer & explanation
Correct answer: B
PK Chunking is a feature specifically designed for this scenario. It automatically splits Bulk API queries on large tables into smaller chunks based on the record's primary key (PK), which is always indexed. This avoids query timeouts and allows for efficient, parallel processing of the chunks without requiring manual calculation of query filters.
- Question 9Intermediate
Data Migration · Deferred Sharing Calculation
During a large data migration of millions of records with complex ownership and sharing rule dependencies, the migration is running extremely slowly. The architect suspects that the constant, synchronous recalculation of sharing rules after each batch is the primary bottleneck.
What is the most effective way to improve performance in this scenario?
Show answer & explanation
Correct answer: B
This is the correct approach. The Deferred Sharing feature is designed for this exact purpose. An administrator can suspend sharing calculations before the migration begins, which prevents the system from re-evaluating access for every record loaded. After the migration is complete, the calculations can be resumed and run as a single, large background process, dramatically speeding up the data load.
- Question 10Beginner
Data Modeling / Database Design · Master-Detail vs. Lookup
An architect is designing a data model for a recruiting application. There will be an
Application__cobject and aReview__cobject. Multiple reviews must be associated with each application. If an application record is deleted, all of its associated review records must also be deleted automatically.Which type of relationship should be created on the
Review__cobject to link it to theApplication__cobject?Show answer & explanation
Correct answer: B
A master-detail relationship is a tightly coupled relationship where the detail (child) record's existence is dependent on the master (parent). This relationship type provides a cascading delete feature, meaning if the master
Application__crecord is deleted, all detailReview__crecords are automatically deleted as well. This perfectly matches the requirement.
Ready for the real thing?
The full DATA-ARCHITECT simulator has every exam-style question, timed mode, and instant scoring.