Business Analytics for Decision Making — Module 4
Course Code: COM1MN110 • Lecture Notes
- Foundational: Taxonomy of Data & Information in Business Analytics In modern enterprise analytics, data constitutes the fundamental raw material from which organizational intelligence and competitive advantage are engineered. However, confusing raw data with refined business information leads to strategic paralysis. A Datum (plural: Data) is a discrete, uncontextualized representation of a fact, transaction, or observation. When data is systematically verified, categorized, structured, and contextualized, it transforms into Information that directly informs managerial choices. 1.1 The Cognitive Value Chain: Data, Information & Business Intelligence Dimension Raw Data Contextualized Information Actionable Business Intelligence Conceptual Definition Unorganized, raw facts, numbers, characters, or signals lacking contextual meaning.
Data that has been processed, structured, cleaned, and organized to answer "Who,
What, Where, When". Synthesized knowledge and predictive insights that answer "Why and What Next" to drive optimal decisions.
Degree of Processing Zero processing; atomic capture from sensors, cash registers, or web clicks.
Aggregated, filtered, formatted into relational tables, reports, or charts.
Statistically modeled, benchmarked against historical baselines, evaluated with scenario simulations.
Managerial Utility Virtually unusable for executive decisionmaking on its own.
Enables operational monitoring, variance tracking, and historical audits.
Guides strategic capital investments, product pivots, and dynamic pricing strategies.
Concrete Example "TXN_90812, 104, 202411-12, 4500, CC_04" "Store #104 generated ₹4,500 in credit card sales on November 12, 2024." "Store #104's credit sales are 25% below the seasonal baseline due to a competitor's festival discount campaign." 1.2 The Seven Inviolable Attributes of High-Value Information
- Accuracy &: Veracity Information must reflect physical reality without distortion, clerical errors, or computational bugs.
Erroneous information leads to the GIGO (Garbage In, Garbage Out) catastrophe.
- Timeliness &: Low Latency Information must reach executive decision-makers before the window of opportunity closes. Stale historical reports lose analytical value in dynamic, real-time markets.
- Relevance &: Actionability Must directly pertain to the specific managerial problem under consideration. An abundance of irrelevant data creates cognitive fatigue and obscures critical trends.
- Completeness &: Sufficiency Must contain all necessary parameters, variables, and dimensions required for comprehensive analysis, without leaving unobserved confounding blind spots. 1.3 Stevens' Four Measurement Scales of Data In 1946, psychologist S.S. Stevens established the foundational taxonomy of data measurement scales, which dictates the mathematical operations and statistical algorithms permissible on any variable:
Measurement Scale Mathematical Characteristics Permissible Mathematical & Statistical Operations Commercial Business Examples
- Nominal: Scale Categorical labels or names used exclusively for identification. No inherent quantitative order, distance, or true zero.
Counting, Frequency Distribution, Mode, Contingency Tables, ChiSquare (χ2) tests. (Cannot compute Mean or Median).
Payment Method (Cash, UPI, Credit Card, Net Banking);
Customer Gender; Retail Store Branch Code; Product Category.
- Ordinal: Scale Categorical data with a clear, meaningful rank order or hierarchy. Intervals between ranks are unequal and mathematically undefined.
Ranking, Median, Percentiles, Interquartile Range (IQR), Spearman's Rank Correlation (ρ), Nonparametric tests.
Customer Satisfaction Ratings (1-Star to 5-Stars); Credit Rating Tiers (AAA, AA, BBB, C);
Employee Performance Ranks; Education Levels.
- Interval: Scale Numerical data with equal, standardized intervals between successive points, but lacks an absolute true zero (zero is arbitrary).
Addition, Subtraction, Arithmetic Mean, Standard Deviation, Variance,
Pearson's Correlation (r), Regression. (Cannot compute ratios).
Temperature in Celsius/Fahrenheit (0°C does not mean absence of heat);
Calendar Year (Year 2024); Credit Bureau Risk Scores (300 to 900).
- Ratio: Scale Continuous numerical data with equal intervals and an absolute, non-arbitrary true physical zero (zero represents complete absence).
All arithmetic operations (Multiplication, Division,
Ratios), Geometric Mean, Coefficient of Variation, Elasticity Modeling.
Sales Revenue in Rupees (₹0 means zero revenue); Order Quantity (units); Customer Age (years); Delivery Lead Time (hours); Distance (km). 1.4 Data Classification by Structural Organization Structured Data (10%–20% of enterprise data): Highly organized data conforming to a rigid predefined schema (schema-on-write). Stored in relational database management systems (RDBMS) as rows and columns. Easily queried via SQL (e.g., accounting general ledgers, inventory stock tables).
- Semi-Structured Data: Lacks a rigid tabular structure but contains internal self-describing markers, tags, or hierarchies (schema-on-read). Examples include JSON payloads from web APIs, XML electronic data interchange (EDI) invoices, and NoSQL key-value stores.
Unstructured Data (80%–90% of enterprise data): Massive volume data lacking any pre-defined data model or schema. Includes customer service phone call audio recordings, executive email text, product video reviews, PDF contracts, and satellite imagery. Processed via Natural Language Processing (NLP) and Deep Learning.
- Secondary: Data: Sources, Strategic Advantages & Critical Hazards Data is broadly classified by its provenance into Primary Data and Secondary Data. Secondary Data refers to information that was previously collected, compiled, organized, and published by an external entity for an objective other than the specific research problem currently under investigation. 2.1 Strategic Advantages & Justification for Secondary Sourcing
- Extreme: Cost Efficiency Acquiring secondary data costs a tiny fraction of primary field research. Government repositories and academic databases are frequently open-access, eliminating the substantial expenses of hiring field interviewers, designing paper forms, and funding travel.
- Immediate: Acquisition Speed Primary nationwide surveys require months of fieldwork, training, and manual coding. Secondary datasets can be downloaded or connected via cloud APIs in minutes, accelerating time-to-insight for rapid strategic decisions.
- Longitudinal: Breadth & Historical Baselines Secondary records provide multi-decade time series trajectories (e.g., national GDP, 30-year inflation, demographic growth) that are physically impossible for a single enterprise to collect retrospectively through primary surveys.
- Feasibility for: Macro-Level Benchmarking Enables corporate strategists to benchmark enterprise market share against industry-wide aggregate production and retail sales figures compiled by trade chambers and government statistical bodies. 2.2 Comprehensive Taxonomy of Secondary Data Sources Source Category Primary Institutional Providers Nature of Information Provided Enterprise Strategic Utility
- National: Government & Regulatory Bodies Census of India, Reserve Bank of India (RBI), MoSPI,
National Sample Survey Office (NSSO), DGCI&S. Macroeconomic indicators, demographic census statistics, national trade balances, agricultural yields.
Macro-environmental scanning, country-level market potential sizing, geographic expansion feasibility.
- International: Multilateral Organizations World Bank Open Data,
International Monetary Fund (IMF), United Nations (UN),
OECD, World Trade Organization (WTO). Global cross-border capital flows, foreign exchange reserves, sovereign debt ratios, global purchasing power parity (PPP).
Cross-border export planning, foreign direct investment (FDI) risk assessment, global supply chain sourcing.
- Commercial: Syndicated Research Firms NielsenIQ, Kantar,
Euromonitor, CRISIL, ICRA, Bloomberg, Refinitiv, Gartner, IDC.
Retail point-of-sale scanner audits, consumer FMCG brand penetration, corporate credit ratings, equity markets.
Competitor market share benchmarking, pricing elasticity tracking, technology stack evaluation.
- Industry: Chambers & Trade Associations Confederation of Indian Industry (CII), FICCI,
ASSOCHAM, NASSCOM, SIAM (Society of Indian Automobile Manufacturers).
Sector-specific production volumes, vehicle registration tallies, tech export revenues, policy whitepapers.
Industry growth forecasting, regulatory advocacy, competitive capacity utilization benchmarking. 2.3 The Inherent Hazards & Operational Pitfalls of Secondary Data Despite its convenience, uncritical reliance on secondary data introduces severe strategic risks into enterprise analytics:
- Temporal Lag & Obsolescence: Official census and government surveys often exhibit publication lags of 2 to 5 years. Utilizing stale demographic data during periods of rapid post-pandemic urbanization leads to severely flawed retail store placement decisions.
Definitional Incompatibility & Unit Mismatch: External agencies often define terms differently than the enterprise. For example, a government agency might define an "active worker" as someone employed for at least 30 days a year, whereas a bank requires continuous 12-month payroll records to evaluate creditworthiness.
Hidden Organizational Bias & Subjective Sponsoring: Research reports published by commercial trade groups or vendor advocacy lobbies frequently cherry-pick findings to promote specific legislative policies or software purchases.
Unknown Sampling Protocols & Non-Response Bias: Analysts rarely have visibility into how field investigators handled non-response, replaced dropouts, or calculated post-stratification weights, potentially masking fatal methodological flaws.
- Inappropriate Levels of Aggregation: Secondary data is frequently aggregated to broad national or state levels, making it impossible to perform micro-level neighborhood or customer-persona analytics. 2.4 The 6-C Evaluation Framework for Secondary Data Auditing
- AUDIT PROTOCOL: THE 6-C EVALUATION FRAMEWORK FOR SECONDARY DATA Quality Assurance CREDIBILITY • CURRENCY • CONSISTENCY • COMPLETENESS • CONTEXT • COVERAGE Critical Evaluation Questions:
1. Credibility: What is the reputation, technical expertise, and institutional independence of the collecting organization?
2. Currency: When was the data physically collected, and does it reflect modern operational realities?
3. Consistency: Do the findings align with independent parallel datasets from other authoritative sources?
4. Completeness: Are key operational variables, confidence intervals, and standard errors fully reported without missing gaps?
5. Context: Under what historical, political, or economic circumstances was the original survey executed?
6. Coverage: Does the sample frame genuinely represent the exact target demographic population of interest?
3. Internal vs. External Enterprise Data Architecture Enterprise data architecture categorizes data pipelines based on organizational boundaries. Modern decision systems synthesize internal operational telemetry with external environmental intelligence to achieve a 360degree view of enterprise performance. 3.1 Internal Data Sourcing: The Digital Footprint of Operations Internal data originates from an organization's daily business operations and represents proprietary, highly detailed records of customer transactions, operational efficiencies, and asset utilization:
- Enterprise: Resource Planning (ERP) Systems Core financial ledgers, bills of materials (BOM), payroll disbursements, supplier purchase orders, and plant production yield logs (SAP, Oracle).
Provides granular records of manufacturing costs and operational throughput.
- Customer: Relationship Management (CRM) Systems Sales pipeline conversion rates, customer interaction histories, support ticket resolution durations, and customer service escalation transcripts (Salesforce, HubSpot). Drives customer churn modeling and sales forecasting.
- Supply: Chain & Logistics Telemetry Radio-frequency identification (RFID) pallet scans,
GPS vehicle fleet route telemetry, warehouse automated storage retrieval system logs, and supplier on-time delivery rates. Optimizes inventory levels and delivery routes.
- Digital: Web & Mobile Clickstream Telemetry Micro-level web server event logs: user session durations, click paths, checkout funnel drop-offs, search queries, and mobile touch heatmaps. Powers personalized algorithmic product recommendation engines. 3.2 External Data Sourcing: Environmental & Competitive Telemetry Internal data alone leaves an enterprise blind to broader market transformations. External data captures macroeconomic, regulatory, and competitive market dynamics:
External Data Stream Primary Technical Sourcing Method Analytical Business Objective Concrete Enterprise Use Case Competitive Web Telemetry & Pricing Automated web scraping, competitor product catalog APIs, digital price aggregators.
Dynamic competitive pricing, product assortment gap analysis, discount monitoring.
An airline scraping competitor ticket prices to adjust dynamic revenue management fares every 10 minutes.
Macroeconomic & Financial Market Feeds Real-time financial exchange feeds (NSE,
BSE, Bloomberg), central bank APIs (RBI). Foreign exchange risk hedging, inflationadjusted cost modeling, interest rate sensitivity.
An import-export firm dynamically managing currency forward contracts based on foreign exchange volatility.
Geospatial & Satellite Imagery Data Commercial satellite feeds, GPS mobile foottraffic aggregators,
Google Maps APIs. Retail store site selection, agricultural commodity harvest forecasting, logistics routing.
A retail chain analyzing parking lot satellite imagery and mobile GPS footfall to predict competitor quarterly earnings.
Social Media & Web Sentiment Streams Twitter/X streaming APIs,
Reddit webhooks, Google Trends, consumer review aggregators.
Brand sentiment tracking, crisis detection, emerging consumer preference discovery.
A consumer electronics brand monitoring real-time sentiment after a smartphone launch to detect battery overheating complaints. 3.3 Enterprise Data Integration: The Master Data Management (MDM) Challenge Merging internal and external data streams introduces severe technical hurdles that require modern Master Data Management (MDM) architectures:
- Entity Resolution & Record Linkage: A single real-world customer may appear as "Rajesh Kumar" in the retail ERP, "R. Kumar" on the mobile app, and "rajesh_k@gmail.com" on social media. Probabilistic recordlinkage algorithms (Levenshtein distance, Jaro-Winkler) must unify these disparate records into a single Golden Record.
- Schema Heterogeneity: Reconciling relational SQL tables with semi-structured JSON external streams using modern Cloud Data Lakehouses (Databricks, Snowflake) that enable unified analytical querying across both paradigms.
- Scientific: Methods of Primary Data Collection When secondary sources are outdated, incomplete, or non-existent, organizations must collect Primary Data —first-hand empirical evidence gathered specifically to address the research problem at hand. Primary data collection utilizes four classical methodologies: Observation, Questionnaires, In-Depth Interviews, and Abstraction from Records. 4.1 Direct Observation & Inspection Methods Observation involves systematically recording the behavioral patterns of people, objects, and events without directly communicating with or questioning the subjects:
1. Human vs. Mechanical Observation
- Human Observation: Trained observers record shopper pathways through retail supermarket aisles or mystery shoppers evaluate customer service protocols.
- Mechanical / Automated Observation: Automated IoT barcode scanners, computerized turnstiles, optical eye-tracking headsets measuring advertising focal points, and in-store computer vision heatmaps.
2. Disguised vs. Undisguised Observation
- Disguised Observation: Subjects are unaware they are being monitored (e.g., hidden security cameras tracking checkout queue times), completely eliminating the Hawthorne Effect (subjects altering behavior when aware of observation).
- Undisguised Observation: Subjects know an observer is present, introducing potential social desirability distortion. 4.2 Survey Questionnaires: Design Architecture & Administration Modes The structured questionnaire is the most widely deployed primary data instrument in commercial marketing and organizational research. Its scientific validity depends heavily on rigorous design principles:
Questionnaire Administration Mode Relative Cost Data Collection Speed Response Rate Primary Risk / Vulnerability Online Web / Mobile Surveys Extremely Low Instantaneous (Days) Low (2%– 10%) Non-response bias; unrepresentative of non-digital demographics.
Telephone Surveys (CATI) Moderate Fast (1–2 Weeks) Moderate (10%–25%) Spam call screening, restricted questionnaire length.
Face-to-Face Personal Surveys Extremely High Slow (Months) High (60%– 80%) Interviewer bias, high labor cost, geographic constraints.
Postal Mail Questionnaires Moderate Extremely Slow Very Low (< 5%) Severe attrition, complete lack of interviewer clarification.
- Questionnaire Design Principles: Avoiding Flawed Formulations
- Eliminating Leading Questions: Questions must be completely neutral. Avoid biased formulations such as "Do you agree that our superior customer service exceeds competitors?" (biased toward agreement). Use: "How would you rate our customer service compared to competitors?" Eliminating Double-Barreled Questions: Never combine two distinct inquiries into a single item, such as "Are you satisfied with our product's price and quality?" (A customer may love the quality but hate the price). Split into two separate questions.
- Funnel Sequencing: Begin with broad, easy, non-threatening opening questions; place complex, cognitive product-evaluation questions in the middle; reserve sensitive demographic questions (income, age) for the very end. 4.3 Scaling Techniques in Survey Research Attitudinal and behavioral variables are quantified using standardized psychometric scales:
- The Likert Scale: The most popular psychometric tool. Presents a declarative statement and asks respondents to indicate their degree of agreement across a balanced 5-point or 7-point continuum (Strongly Disagree, Disagree, Neutral, Agree, Strongly Agree).
- The Semantic Differential Scale: Employs a 7-point scale anchored at either end by bipolar opposite adjectives (e.g., Modern [1 - 2 - 3 - 4 - 5 - 6 - 7] Obsolete; Reliable [1 - 2 - 3 - 4 - 5 - 6 - 7] Unreliable). 4.4 In-Depth Interviews & Focus Group Discussions (FGD)
1. Structured vs. Semi-Structured Interviews
- Structured Interviews: Rigid adherence to a predetermined interview schedule without deviation; maximizes quantitative comparability across respondents.
- Semi-Structured / In-Depth Interviews (IDI): Guided by an open-ended discussion framework allowing skilled interviewers to probe underlying emotional drivers, subconscious motivations, and complex B2B decision journeys.
- Focus: Group Discussions (FGD) A moderated qualitative discussion among 6 to 10 prescreened target participants lasting 90 to 120 minutes. Leverages group interaction to generate rich exploratory ideas for new product packaging, branding concepts, and advertising storyboards. 4.5 Abstraction from Records & Published Statistics Data Abstraction involves systematically extracting quantitative variables from existing physical or electronic operational records—such as auditing 5,000 corporate legal contracts, hospital patient medical discharge summaries, or warranty repair claims. It requires rigorous coding rubrics and dual-auditor inter-rater reliability checks to ensure objective data extraction.
- Data: Quality Management, Hygiene & Worked Quantitative Illustrations To convert raw collected data into reliable enterprise analytics, data scientists must ensure instrument reliability, determine statistically representative sample sizes, audit data quality, and comply with modern data privacy statutes. 5.1 Instrument Reliability: Mathematical Formulation of Cronbach's Alpha When survey questionnaires measure latent constructs (e.g., customer brand loyalty or employee burnout) using multiple Likert items, analysts must verify internal consistency reliability via Cronbach's Alpha (α):
- MATHEMATICAL FORMULATION: CRONBACH'S ALPHA INTERNAL CONSISTENCY RELIABILITY Psychometric Quality Audit α = [ k ÷ (k − 1) ] × [ 1 − ( ∑ σ_i2 ÷ σ_total2 ) ] Where: k: Total number of survey scale items (questions measuring the single latent construct). ∑ σ_i2: Sum of the individual variances of each of the k survey items. σ_total2: Variance of the total composite test scores across all respondents.
- Benchmark Standards: α ≥ 0.90 (Excellent); 0.80 ≤ α < 0.90 (Good); 0.70 ≤ α < 0.80 (Acceptable); α < 0.70 (Poor/Unreliable scale requiring item deletion). 5.2 Sample Size Determination in Survey Research Collecting too few respondents introduces fatal sampling error; surveying too many wastes financial resources. The optimal scientific sample size for estimating a population proportion is calculated using Cochran's classic formula:
- MATHEMATICAL FORMULATION: SAMPLE SIZE FOR ESTIMATING PROPORTIONS Sampling Theory Infinite Population: n_0 = ( Z2 × p × (1 − p) ) ÷ E2 Finite Population Correction: n = n_0 ÷ [ 1 + ( (n_0 − 1) ÷ N ) ] Where:
Z: Critical value from standard normal distribution at desired confidence level (Z = 1.96 for 95% confidence; Z = 2.576 for 99% confidence). p: Estimated baseline proportion exhibiting the characteristic (use p = 0.50 for maximum variance conservatism if unknown).
E: Acceptable margin of error expressed as a decimal (e.g., ±3% error → E = 0.03).
N: Total size of the finite target population. 5.3 Worked Quantitative Illustrations in Data & Information Analytics ∑ Worked Illustration 1: Questionnaire Internal Consistency Reliability (Cronbach's Alpha) Customer Mobile App Usability Survey Scale Audit:
- Scale consists of k = 4 Likert items (rated 1 to 5) administered to a pilot test sample.
- Calculated Individual Item Variances:
- Item 1 Variance σ_12 = 0.80
- Item 2 Variance σ_22 = 0.65
- Item 3 Variance σ_32 = 0.75
- Item 4 Variance σ_42 = 0.80 → Sum of Item Variances ∑σ_i2 = 0.80 + 0.65 + 0.75 + 0.80 = 3.00.
- Variance of Total Composite Composite Scores across all respondents: σ_total2 = 8.50.
- Correction: Factor Calculation: → k ÷ (k − 1) = 4 ÷ (4 − 1) = 4 ÷ 3 = 1.3333.
- Variance: Ratio Calculation: → ∑σ_i2 ÷ σ_total2 = 3.00 ÷ 8.50 = 0.3529. → [ 1 − (∑σ_i2 ÷ σ_total2) ] = 1 − 0.3529 = 0.6471.
- Final: Cronbach's Alpha Computation: → α = 1.3333 × 0.6471 = 0.8628 ≈ 0.863.
- PSYCHOMETRIC RELIABILITY VERDICT: With α = 0.863 (> 0.80 threshold), the 4-item usability survey demonstrates strong internal consistency reliability, validating its deployment for enterprise-wide customer surveys. ∑ Worked Illustration 2: Determining Optimal Sample Size for Retail Customer Research Retail Chain Customer Satisfaction Survey Parameters:
- Desired Confidence Level: 95% → Critical Normal Value Z = 1.96 (Z2 = 3.8416).
- Acceptable Margin of Error: E = ±4.0% (0.04 → E2 = 0.0016).
- Expected Population Proportion: Conservative default p = 0.50 (p(1−p) = 0.50 × 0.50 = 0.25).
- Total Registered Loyalty Customer Population: N = 5,000 Customers (Finite Population).
- Initial: Sample Size for Infinite Population (n_0): → n_0 = (Z2 × p × (1 − p)) ÷ E2 → n_0 = (3.8416 × 0.25) ÷ 0.0016 = 0.9604 ÷ 0.0016 = 600.25 Respondents.
- Finite: Population Correction Adjustment: → n = n_0 ÷ [ 1 + ( (n_0 − 1) ÷ N ) ] → (n_0 − 1) ÷ N = (600.25 − 1) ÷ 5,000 = 599.25 ÷ 5,000 = 0.11985. → Denominator = 1 + 0.11985 = 1.11985. → n = 600.25 ÷ 1.11985 = 536.01 ≈ 537 Completed Surveys.
- SAMPLING DESIGN VERDICT: The research team must sample exactly 537 verified loyalty members to guarantee statistical estimates within ±4% margin of error at a 95% confidence level, saving the firm the expense of over-surveying. ∑ Worked Illustration 3: Quantitative 6-C Audit of Commercial Secondary Data Providers Procurement Evaluation of Two Commercial Syndicated Datasets:
- Provider A (Established Global Research House, Annual Fee ₹25 Lakhs)
- Provider B (Niche Web Analytics Startup, Annual Fee ₹8 Lakhs)
- Weighted 6-C Audit Criteria (Weights sum to 1.00):
- Credibility &: Reputation (w_1 = 0.25)
- Currency &: Update Frequency (w_2 = 0.20)
- Consistency with: Parallel Data (w_3 = 0.15)
- Completeness &: Granularity (w_4 = 0.20)
- Coverage of: Target Demographics (w_5 = 0.20)
- Audit: Scores (Rated 1 to 10 Scale by Chief Data Officer):
- Provider A: Credibility = 9.5 | Currency = 6.0 (Quarterly lag) | Consistency = 9.0 | Completeness = 8.5 | Coverage = 8.0
- Provider B: Credibility = 6.0 | Currency = 9.5 (Real-time feeds) | Consistency = 7.0 | Completeness = 7.5 | Coverage = 6.5
- Weighted: Audit Score Calculation [Score = ∑ w_i × Rating_i]:
- Provider A: (0.25 × 9.5) + (0.20 × 6.0) + (0.15 × 9.0) + (0.20 × 8.5) + (0.20 × 8.0) → Score A = 2.375 + 1.20 + 1.35 + 1.70 + 1.60 = 8.225 / 10.0
- Provider B: (0.25 × 6.0) + (0.20 × 9.5) + (0.15 × 7.0) + (0.20 × 7.5) + (0.20 × 6.5) → Score B = 1.50 + 1.90 + 1.05 + 1.50 + 1.30 = 7.250 / 10.0
- PROCUREMENT DECISION: Provider A achieves a superior audit score of 8.225, justifying its premium fee due to higher credibility, demographic coverage, and methodological consistency essential for executive board reporting. 5.4 Data Ethics, Privacy & Legal Frameworks in Data Collection In the era of Big Data, data collection is strictly bounded by legal and ethical frameworks:
- Principle of Informed Consent: Survey respondents and digital app users must be explicitly informed of what data is collected, how it will be processed, and for what specific commercial purposes. Consent must be freely given, specific, and revocable.
- Data Minimization: Enterprises must collect only the minimum amount of personal data strictly necessary to fulfill the immediate analytical objective (e.g., an e-commerce checkout does not require collecting an applicant's political views or health records).
- Statutory Compliance: Strict adherence to the Digital Personal Data Protection Act (DPDP 2023) of India and the European Union General Data Protection Regulation (GDPR), which enforce massive statutory penalties (up to ₹250 Crore in India; €20 Million or 4% of global turnover under GDPR) for unauthorized personal data collection, unauthorized secondary use, or inadequate security safeguarding. 5.5 Master Analytical Synthesis: Data Collection Strategy Matrix Research Objective / Problem Context Optimal Data Source Primary Collection Methodology Key Quality Safeguard Macroeconomic Country Expansion Secondary External World Bank, RBI, Census published statistics.
Apply the 6-C framework; check currency and reporting lag.
Optimizing In-Store Checkout Queues Primary Internal Automated mechanical video observation, POS logs.
Disguised observation to eliminate the Hawthorne effect.
Measuring Customer Brand Loyalty Primary External Structured online questionnaire with 5-point Likert scale.
Test internal consistency via Cronbach's Alpha (α ≥ 0.80).
Investigating Complex B2B Buying Journeys Primary Qualitative Semi-structured in-depth executive interviews.
Dual-analyst inter-coder reliability checks on interview transcripts.
Detecting Real-Time Banking Fraud Primary & Secondary Internal Real-time streaming transaction logs matched against historical baseline profiles.
Sub-second automated anomaly detection pipelines.
Download Module 4 Notes (PDF)
Calicut University • FYUGP 2024 Syllabus
Finished this module?
Continue reading the next module or return to the subject overview.