the set query benchmark patrick e. o'neil

42
The Set Query Benchmark Patrick E. O’Neil Department of Mathematics and Computer Science University of Massachusetts at Boston Hak Chan Lee School of Computing, Soong Sil Univ. [email protected]

Upload: tess98

Post on 20-Jan-2015

405 views

Category:

Documents


3 download

DESCRIPTION

 

TRANSCRIPT

Page 1: The Set Query Benchmark Patrick E. O'Neil

The Set Query BenchmarkPatrick E. O’Neil

Department of Mathematics and Computer ScienceUniversity of Massachusetts at Boston

Hak Chan LeeSchool of Computing, Soong Sil Univ.

[email protected]

Page 2: The Set Query Benchmark Patrick E. O'Neil

Table of Contents

• Introduction to the benchmark• An application of the benchmark• How to run the set query benchmark

Page 3: The Set Query Benchmark Patrick E. O'Neil

Introduction to the benchmark

• Strategic value data applications– Marketing Information System– Decision Support System– Management Reporting– Direct Marketing

• Set query Benchmark– To aid decision makers who require performance data

relevant to strategic value data applications.• Database

– IBM’s DB2– Computer Corporation of America’s MODEL 204

• Performance– I/O, CPU time, elapsed time

Page 4: The Set Query Benchmark Patrick E. O'Neil

Features of The Set Query Benchmark

• Four key characteristics of Set Query– Portability

• The queries for benchmark are specified in SQL which is available on most systems.

– Functional Coverage• Prospective users should be able a rating for the particular

subset they expect to use most in the system they are planning

– Selectivity Coverage• Applies to a clause of a query, and indicates the proportion of

rows of the database to be selected.• A single row has very high selectivity

– Scalability• The database has a single table, known as the BENCH table,

which contains an integer multiple of 1 million rows of 200bytes each.

Page 5: The Set Query Benchmark Patrick E. O'Neil

Definition of the BENCH Table

• BENCH table– 13 indexed columns (length 4 each)– Character columns

• s1 (length 8)• s2 through s8 (length 20 each)• Never used in retrieval queries

– (13*4) + (1*8) + (7*20) = 200(bytes)– Cardinality (number of distinct value)– Where a BENCH table of more than 1 million

rows is created, the three highest cardinality columns are either renamed or reinterpreted in a consistent way

Page 6: The Set Query Benchmark Patrick E. O'Neil

First 10 rows of the BENCH database

Page 7: The Set Query Benchmark Patrick E. O'Neil

Achieving Functional Coverage

• Five companies with state of the art strategic information applications were contacted

• Three general strategic value data applications– Document search– Direct marketing– Decision support / management reporting

Page 8: The Set Query Benchmark Patrick E. O'Neil

Document search• The “documents” can represent any information

• (i) A COUNT of records with a single exact match condition, known as query Q1:

Q1: SELECT COUNT (*) FROM BENCHWHERE KN=2;

For each KN ε {KSEQ, K100K,…, K4, K2}

• (ii) A COUNT of records from a conjunction of two exact match condition: query Q2A:

Q2A: SELECT COUNT (*) FROM BENCHWHERE K2 = 2 AND KN = 3;

For each KN ε {KSEQ, K100K,…, K4}Or

Q2B: SELECT COUNT (*) FROM BENCHWHERE K2 = 2 AND NOT KN = 3;

For each KN ε {KSEQ, K100K,…, K4}

Page 9: The Set Query Benchmark Patrick E. O'Neil

Document search (cont.)

• (iii) A retrieval of data (not counts) given constraints of three conditions, including range conditions, (Q4A), or constraints of five conditions, (Q4B).

Q4: SELECT KSEQ, K500K FROM BENCHWHERE constraint with (3 or 5) conditions;

Page 10: The Set Query Benchmark Patrick E. O'Neil

Direct Marketing

• Goal– To identify a list of households which are most

likely to purchase a given product or service

• Approach selecting such a list– Preliminary sizing and exploration of possible

criteria for selection• R.L. Polk • reinforced Q2A and Q2B

– Retrieving the data from the records for a mailing or other communication

• Saks Fifth Avenue• reinforced Q4A and Q4B

Page 11: The Set Query Benchmark Patrick E. O'Neil

Direct Marketing (cont.)Q3A: SELECT SUM (K1K) FROM BENCH

WHERE KSEQ BETWEEN 400000 AND 500000 AND KN = 3;For each KN ε {K100K,…, K4}

Q3B: SELECT SUM (K1K) FROM BENCHWHERE (KSEQ BETWEEN 400000 AND 410000OR KSEQ BETWEEN 420000 AND 430000OR KSEQ BETWEEN 440000 AND 450000OR KSEQ BETWEEN 460000 AND 470000OR KSEQ BETWEEN 480000 AND 500000)

AND KN = 3; For each KN ε {K100K,…, K4}

Page 12: The Set Query Benchmark Patrick E. O'Neil

Decision Support and Management Reporting

Q5: SELECT KN1, KN2, COUNT (*) FROM BENCHGROUP BY KN1, KN2;

For each(KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Q6A: SELECT COUNT (*) FROM BENCH B1, BENCH B2WHERE B1.KN = 49 AND B1.K250K = B2.K500K;

For each (KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Q6B: SELECT B1.KSEQ, B2.KSEQ FROM BENCH B1, BENCH B2WHERE B1.KN = 99AND B1.K250K = B2.K500KAND B2.K25 = 19;

For each (KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Page 13: The Set Query Benchmark Patrick E. O'Neil

Running the Benchmark

• Architectural Complexity• A Single Unifying Criterion• Confidence in the Result• Output Formatting

Page 14: The Set Query Benchmark Patrick E. O'Neil

Architectural Complexity

• Measuring these two products exposes an enormous number of architectural features peculiar to each.

• Perfect apples-to-apples comparison between two products in every features is generally impossible.

• MODEL 204– Uses standard (random) I/O– Existence Bit Map

• Is able to perform much more efficient negation queries.

• DB2– Uses Random I/O and Prefetch I/O

Page 15: The Set Query Benchmark Patrick E. O'Neil

A Single Unifying Criterion

• Sum up our benchmark results with a single figure– Dollar Price per Query Per Second

($PRICE/QPS)

Page 16: The Set Query Benchmark Patrick E. O'Neil

Confidence in the Result

• Table 2.2.1, containing Q1measurements, displays a maximum elapsed time of 39.69 seconds in one case and 0.31seconds in the other

• How can we have confidence that we have not made a measurement or tuning mistake when confronted with variation of this kind?– The answer lies in understanding the two system

architectures.

Page 17: The Set Query Benchmark Patrick E. O'Neil

Output Formatting

• The query results are directed to a user file which represents the answers printable form. – About half the reports return very few records (only

counts)– About half return a few thousand records.

Page 18: The Set Query Benchmark Patrick E. O'Neil

An Application of the Benchmark

• Hardware/Software environment– DB2

• Version 2.2 with 1200 4KByte memory buffer

– MODEL 204• Version 2.1 with 800 6KByte buffer

– 3380 disk drive

– Standalone 4381-3, a 4.5 MIPS dual processor.

– OS• MVS XA 2.2 Operating System

Page 19: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered

• RUNSTATS– To update the various tables, such as

SYSCOLUMNS and SYSINDEXES– Examine the data– Accumulate statistics

Page 20: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• The rows of the BENCH database were loaded in the DB2 and MODEL 204 database in identical fashion.

• Indices in both cases were loaded with leaf nodes 95% full

Page 21: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• MODEL 204– Lager pages (6Kbytes)– B-tree indices have complex indirect inversion structures– B-tree entries need not contain lists of RIDs of rows with duplicat

e values.– But may point to a list or “bitmap” of such RIDs.– more efficient

Page 22: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• In MODEL 204– The native “User language” was used to

formulates these queries, rather than SQL.

Page 23: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q1: Count, Single Exact MatchSELECT COUNT (*) FROM BENCHWHERE KN=2;

For each KN ε {KSEQ, K100K,…, K4, K2}

Page 24: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• DB2 times com from DB2PM Class 1• Class 1 statistics

– Include time spent by SPUFI application and cross-region communication

• Class 2 statistics– Include only time spent within DB2

• Total number of pages (N)– R : random reads– P : prefetch reads– N = R + 32 * P

• DB2 slowdown in not caused by the I/O

Page 25: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q2: Count ANDing two ClausesQ2A: SELECT COUNT (*) FROM BENCH

WHERE K2 = 2 AND KN = 3;For each KN ε {KSEQ, K100K,…, K4}

Page 26: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

Q2B: SELECT COUNT (*) FROM BENCHWHERE K2 = 2 AND NOT KN = 3;

For each KN ε {KSEQ, K100K,…, K4}

Page 27: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q3: Sum, Range and Match ClauseQ3A: SELECT SUM (K1K) FROM BENCH

WHERE KSEQ BETWEEN 400000 AND 500000 AND KN = 3;For each KN ε {K100K,…, K4}

Page 28: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)Q3B: SELECT SUM (K1K) FROM BENCH

WHERE (KSEQ BETWEEN 400000 AND 410000OR KSEQ BETWEEN 420000 AND 430000OR KSEQ BETWEEN 440000 AND 450000OR KSEQ BETWEEN 460000 AND 470000OR KSEQ BETWEEN 480000 AND 500000)

AND KN = 3; For each KN ε {K100K,…, K4}

Page 29: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q4: Multiple Condition SelectionQ4: SELECT KSEQ, K500K FROM BENCH

WHERE constraint with (3 or 5) conditions;

• Condition Sequence– (1) K2 = 1; (2) K100 > 80;– (3) K100K BETWEEN 2000 and 3000; (4) K5 = 3;– (5) K25 IN (11, 19); (6) K4 = 3;– (7) K100 < 41; (8) K1K BETWEEN 850 AND

950;– (9) K10 = 7; (10) K25 IN (3, 4);

Page 30: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

Page 31: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q5: Multiple Column GROUP BYQ5: SELECT KN1, KN2, COUNT (*) FROM BENCH

GROUP BY KN1, KN2;For each(KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Page 32: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)

• Q6: Join ConditionQ6A: SELECT COUNT (*) FROM BENCH B1, BENCH B2

WHERE B1.KN = 49 AND B1.K250K = B2.K500K; For each

(KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Page 33: The Set Query Benchmark Patrick E. O'Neil

Statistics Gathered (cont.)Q6B: SELECT B1.KSEQ, B2.KSEQ FROM BENCH B1, BENCH B2

WHERE B1.KN = 99AND B1.K250K = B2.K500KAND B2.K25 = 19;

For each(KN1, KN2) ε {(K2, K100), (K10, K25), (K10, K25)}

Page 34: The Set Query Benchmark Patrick E. O'Neil

Single Measure Result

• The ultimate measure of Set Query benchmark for any platform measured is price per query per second ($/QPS).

Page 35: The Set Query Benchmark Patrick E. O'Neil

How to Run the Set Query Benchmark

• Generating the Data• Running the Queries• Interpreting the Result• Configuration and Reporting

Requirement – a checklist• Concluding Remark

Page 36: The Set Query Benchmark Patrick E. O'Neil

Generating the Data

• Pseudo-random number generator is suggested by Ward Cheney and David Kincaid

Page 37: The Set Query Benchmark Patrick E. O'Neil

Generating the Data (cont.)

• Loading the Data– Indexes for 13 columns– Indices

• B-tree type (if available)• Loaded 95% full

– Set Query benchmark permit the data to be loaded on as many disk devices as desired.

Page 38: The Set Query Benchmark Patrick E. O'Neil

Running the Queries

• To support tuning decisions regarding memory purchase, the Set Query benchmark offers two approaches– It is possible to scale the BENCH table to

any integer multiple of one million bytes.– The benchmark measurements are

normalized by essentially flushing the buffers between successive queries, so that the number of I/Os is maximum

Page 39: The Set Query Benchmark Patrick E. O'Neil

Interpreting the result

• A common demand of decision maker who require benchmark result is a single number.

• Calculating the $/QPS Rating– The calculation of the $PRICE/QPS rating for the

Set Query benchmark is quite straightforward in concept

• 69 queries• T : Elapsed time to run all of them in series

– 69/T queries per second• P : well defined dollar, for hardware and software• $PRICE/QPS : (P*T)/69

Page 40: The Set Query Benchmark Patrick E. O'Neil

Interpreting the result (cont.)

• Price calculation starts by summing the CPU and I/O resources

• To calculate TF=SUMELA/SUMCPU

T=F*TOTCPU

Page 41: The Set Query Benchmark Patrick E. O'Neil

Configuration and Reporting Requirements – a Checklist

• The data should be generated in the precise manner

• The BENCH table should be loaded on the appropriate number of stable devices, and resource utilization for the load, together with any preparatory utilities to improve performance should be reported

• The exact hardware/software configuration used for test should be reported

• The access strategy for each Query reported should be explained in the report if this capability is offered by the database system.

• The benchmark should be run from either an interactive or an embedded program and the Page Buffer and Sort Buffer (if separate) space in Mbyte should be reported.

• All resource use should be reported for each of the queries.

Page 42: The Set Query Benchmark Patrick E. O'Neil

Concluding Remarks

• Variation in resource use translate into large dollar price difference

• Extensions to the Set Query benchmark have been suggested by a number of observers, and a multi-user Set Query benchmark is under study.