hortonworksodbc driverwithsql connectorforapache spark · pdf filehortonworksodbc...

87
Hortonworks ODBC Driver with SQL Connector for Apache Spark User Guide Revised: December 21, 2015 © 2012-2015 Hortonworks Inc. All Rights Reserved. Parts of this Program and Documentation include proprietarysoftware and content that is copyrighted and licensed by Simba TechnologiesIncorporated. This proprietary software and content may include one or more feature, functionality or methodologywithin the ODBC, JDBC, ADO.NET, OLE DB, ODBO, XMLA, SQL and/or MDX component(s). For information about Simba's products and services, visit: www.simba.com Architecting the Future of Big Data

Upload: dangxuyen

Post on 06-Feb-2018

239 views

Category:

Documents


2 download

TRANSCRIPT

Hortonworks ODBCDriver with SQLConnector for ApacheSparkUser Guide

Revised: December 21, 2015

© 2012-2015 Hortonworks Inc. All Rights Reserved.

Parts of this Program and Documentation include proprietary software and content that is copyrighted and licensed by SimbaTechnologies Incorporated. This proprietary software and content may include one or more feature, functionality ormethodologywithin the ODBC, JDBC, ADO.NET, OLE DB, ODBO, XMLA, SQL and/or MDX component(s).

For information about Simba's products and services, visit: www.simba.com

Architecting the Future of Big Data

Hortonworks Inc. Page 2

Table of Contents

Introduction 4Contact Us 5Windows Driver 6

Windows System Requirements 6Installing the Driver 6Creating a Data Source Name 7Configuring a DSN-less Connection 9Configuring Authentication 10Configuring Advanced Options 15Configuring Server-Side Properties 17Configuring HTTP Options 17Configuring SSL Verification 18Configuring Logging Options 19Configuring Kerberos Authentication for Windows 21Verifying the Version Number 25

Linux Driver 26Linux System Requirements 26Installing the Driver 26Setting the LD_LIBRARY_PATH Environment Variable 27Verifying the Version Number 28

Mac OS X Driver 29Mac OS X System Requirements 29Installing the Driver 29Setting the DYLD_LIBRARY_PATH Environment Variable 30Verifying the Version Number 30

Configuring ODBC Connections for Non-Windows Platforms 31Configuration Files 31Sample Configuration Files 32Configuring the Environment 32

Architecting the Future of Big Data

Hortonworks Inc. Page 3

Defining DSNs in odbc.ini 33Specifying ODBC drivers in odbcinst.ini 34Configuring Driver Settings in hortonworks.sparkodbc.ini 35Configuring Authentication 36Configuring SSL Verification 39Testing the Connection 40Configuring Logging Options 42

Authentication Options 44Using a Connection String 46

DSN Connection String Example 46DSN-less Connection String Examples 46

Features 51SQL Connector for HiveQL 51Data Types 51Catalog and Schema Support 52spark_system Table 53Server-Side Properties 53Get Tables With Query 53Active Directory 54Write-back 54

Driver Configuration Options 55Configuration Options Appearing in the User Interface 55Configuration Options Having Only Key Names 70

Contact Us 73Third Party Trademarks and Licenses 74

Architecting the Future of Big Data

Hortonworks Inc. Page 4

IntroductionThe Hortonworks ODBC Driver with SQL Connector for Apache Spark is used for directSQL and HiveQL access to Apache Hadoop / Spark distributions, enabling BusinessIntelligence (BI), analytics, and reporting on Hadoop-based data. The driver efficientlytransforms an application’s SQL query into the equivalent form in HiveQL, which is a subsetof SQL-92. If an application is Spark-aware, then the driver is configurable to pass the querythrough to the database for processing. The driver interrogates Spark to obtain schemainformation to present to a SQL-based application. Queries, including joins, are translatedfromSQL to HiveQL. For more information about the differences between HiveQL and SQL,see "Features" on page 51.

The Hortonworks ODBC Driver with SQL Connector for Apache Spark complies with theODBC 3.80 data standard and adds important functionality such as Unicode and 32- and 64-bit support for high-performance computing environments.

ODBC is one of the most established and widely supported APIs for connecting to andworking with databases. At the heart of the technology is the ODBC driver, which connectsan application to the database. For more information about ODBC, seehttp://www.simba.com/resources/data-access-standards-library. For complete informationabout the ODBC specification, see the ODBC API Reference athttp://msdn.microsoft.com/en-us/library/windows/desktop/ms714562(v=vs.85).aspx.

TheUser Guide is suitable for users who are looking to access data residing within Hadoopfrom their desktop environment. Application developers may also find the informationhelpful. Refer to your application for details on connecting via ODBC.

Architecting the Future of Big Data

Hortonworks Inc. Page 5

Contact UsIf you have difficulty using the Hortonworks ODBC Driver with SQL Connector for ApacheSpark, please contact our support staff. We welcome your questions, comments, andfeature requests.

Please have a detailed summary of the client and server environment (OS version, patchlevel, Hadoop distribution version, Spark version, configuration, etc.) ready, before you callor write us. Supplying this information accelerates support.

By telephone:

USA: (855) 8-HORTON

International: (408) 916-4121

On the Internet:

Visit us at www.hortonworks.com

Architecting the Future of Big Data

Hortonworks Inc. Page 6

Windows DriverWindows System RequirementsYou install the Hortonworks ODBC Driver with SQL Connector for Apache Spark on clientmachines that access data stored in a Hadoop cluster with the Spark service installed andrunning. Each machine that you install the driver on must meet the followingminimumsystem requirements:

l One of the following operating systems:o Windows 7 SP1, 8, or 8.1o Windows Server 2008 R2 SP1, 2012, or 2012 R2

l 100 MB of available disk space

Important: To install the driver, you must have Administrator privileges on thecomputer.

The driver supports Apache Spark versions 0.8 through 1.5.

Installing the DriverOn 64-bit Windows operating systems, you can execute both 32- and 64-bit applications.However, 64-bit applicationsmust use 64-bit drivers and 32-bit applicationsmust use 32-bitdrivers. Make sure that you use the version of the driver matching the bitness of the clientapplication accessing data in Hadoop / Spark:

l HortonworksSparkODBC32.msi for 32-bit applicationsl HortonworksSparkODBC64.msi for 64-bit applications

You can install both versions of the driver on the same computer.

To install the Hortonworks ODBC Driver with SQL Connector for Apache Spark:1. Depending on the bitness of your client application, double-click to run

HortonworksSparkODBC32.msi or HortonworksSparkODBC64.msi.2. ClickNext.3. Select the check box to accept the terms of the License Agreement if you agree, and

then click Next.4. To change the installation location, click Change, then browse to the desired folder,

and then clickOK. To accept the installation location, clickNext.5. Click Install.6. When the installation completes, click Finish.

Architecting the Future of Big Data

Hortonworks Inc. Page 7

Creating a Data Source NameTypically, after installing the Hortonworks ODBC Driver with SQL Connector for ApacheSpark, you need to create a Data Source Name (DSN).

Alternatively, for information about DSN-less connections, see "Configuring a DSN-lessConnection" on page 9.

To create a Data Source Name:1. Open the ODBC Administrator:

l If you are using Windows 7 or earlier, click the Start button , then click AllPrograms, then click the Hortonworks Spark ODBC Driver 1.1 programgroup corresponding to the bitness of the client application accessing data inHadoop / Spark, and then clickODBC Administrator.

l Or, if you are using Windows 8 or later, on the Start screen, type ODBC admin-istrator, and then click theODBC Administrator search result correspondingto the bitness of the client application accessing data in Hadoop / Spark.

2. In the ODBC Data Source Administrator, click the Drivers tab, and then scroll downas needed to confirm that the Hortonworks Spark ODBC Driver appears in thealphabetical list of ODBC drivers that are installed on your system.

3. Choose one:l To create a DSN that only the user currently logged into Windows can use, clicktheUser DSN tab.

l Or, to create a DSN that all users who log into Windows can use, click the Sys-tem DSN tab.

4. ClickAdd.5. In the Create New Data Source dialog box, select Hortonworks Spark ODBC

Driver and then click Finish.6. In theData Source Name field, type a name for your DSN.7. Optionally, in the Description field, type relevant details about the DSN.8. In theHost field, type the IP address or host name of the Spark server.9. In the Port field, type the number of the TCP port that the Spark server uses to listen

for client connections.10. In theDatabase field, type the name of the database schema to use when a schema is

not explicitly specified in a query.Note: You can still issue queries on other schemas by explicitly specifying theschema in the query. To inspect your databases and determine the appropriateschema to use, type the show databases command at the Spark commandprompt.

11. In the Spark Server Type list, select the appropriate server type for the version ofSpark that you are running:

Architecting the Future of Big Data

Hortonworks Inc. Page 8

l If you are running Shark 0.8.1 or earlier, then select SharkServerl If you are running Shark 0.9.*, then select SharkServer2l If you are running Spark 1.1 or later, then select SparkThriftServer

12. In the Authentication area, configure authentication as needed. For more information,see "Configuring Authentication" on page 10.Note: Shark Server does not support authentication. Most default configurations ofShark Server 2 or Spark Thrift Server require User Name authentication. To verifythe authentication mechanism that you need to use for your connection, check theconfiguration of your Hadoop / Spark distribution. For more information, see"Authentication Options" on page 44.

13.Optionally, if the operations against Spark are to be done on behalf of a user that isdifferent than the authenticated user for the connection, type the name of the user tobe delegated in theDelegation UID field.Note: This option is applicable onlywhen connecting to a Shark Server 2 or SparkThrift Server instance that supports this feature.

14. In the Thrift Transport list, select the transport protocol to use in the Thrift layer.15. If the Thrift Transport option is set to HTTP, then to configure HTTP options such as

custom headers, clickHTTP Options. For more information, see "ConfiguringHTTP Options" on page 17.

16.To configure client-server verification over SSL, click SSL Options. For moreinformation, see "Configuring SSL Verification" on page 18.

17.To configure advanced driver options, clickAdvanced Options. For moreinformation, see "Configuring Advanced Options" on page 15.

18.To configure server-side properties, click Advanced Options and then click ServerSide Properties. For more information, see "Configuring Server-Side Properties" onpage 17.Important: When connecting to Spark 0.14 or later, the Temporary Tables featureis always enabled and you do not need to configure it in the driver.

19.To configure logging behavior for the driver, click Logging Options. For moreinformation, see "Configuring Logging Options" on page 19.

20.To test the connection, click Test. Review the results as needed, and then clickOK.Note: If the connection fails, then confirm that the settings in the HortonworksSpark ODBC Driver DSN Setup dialog box are correct. Contact your Spark serveradministrator as needed.

21.To save your settings and close the Hortonworks Spark ODBC Driver DSN Setupdialog box, clickOK.

22. To close the ODBC Data Source Administrator, click OK.

Architecting the Future of Big Data

Hortonworks Inc. Page 9

Configuring a DSN-less ConnectionSome client applications provide support for connecting to a data source using a driverwithout a Data Source Name (DSN). To configure a DSN-less connection, you can use aconnection string or the Hortonworks Spark ODBC Driver Configuration tool that is installedwith the Hortonworks ODBC Driver with SQL Connector for Apache Spark. The followingsection explains how to use the driver configuration tool. For information about usingconnection strings, see "DSN-less Connection String Examples" on page 46.

To configure a DSN-less connection using the driver configuration tool:1. Choose one:

l If you are using Windows 7 or earlier, click the Start button , then click AllPrograms, and then click theHortonworks Spark ODBC Driver 1.1 programgroup corresponding to the bitness of the client application accessing data inHadoop / Spark.

l Or, if you are using Windows 8 or later, click the arrow button at the bottom ofthe Start screen, and then find theHortonworks Spark ODBC Driver 1.1 pro-gramgroup corresponding to the bitness of the client application accessing datain Hadoop / Spark.

2. ClickDriver Configuration, and then click OK if prompted for administratorpermission to make modifications to the computer.Note: Youmust have administrator access to the computer to run this applicationbecause it makes changes to the registry.

3. In the Spark Server Type list, select the appropriate server type for the version ofSpark that you are running:

l If you are running Shark 0.8.1 or earlier, then select SharkServer.l If you are running Shark 0.9.*, then select SharkServer2.l If you are running Spark 1.1 or later, then select SparkThriftServer.

4. In the Authentication area, configure authentication as needed. For more information,see "Configuring Authentication" on page 10.Note: Spark Server 1 does not support authentication. Most default configurationsof Spark Server 2 require User Name authentication. To verify the authenticationmechanism that you need to use for your connection, check the configuration ofyour Hadoop / Spark distribution. For more information, see "AuthenticationOptions" on page 44.

5. Optionally, if the operations against Spark are to be done on behalf of a user that isdifferent than the authenticated user for the connection, then in the DelegationUID field, type the name of the user to be delegated.Note: This option is applicable onlywhen connecting to a Shark Server 2 or SparkThrift Server instance that supports this feature.

6. In the Thrift Transport list, select the transport protocol to use in the Thrift layer.

Architecting the Future of Big Data

Hortonworks Inc. Page 10

7. If the Thrift Transport option is set to HTTP, then to configure HTTP options such ascustom headers, clickHTTP Options. For more information, see "ConfiguringHTTP Options" on page 17.

8. To configure client-server verification over SSL, click SSL Options. For moreinformation, see "Configuring SSL Verification" on page 18.

9. To configure advanced options, clickAdvanced Options. For more information, see"Configuring Advanced Options" on page 15.

10.To configure server-side properties, click Advanced Options and then click ServerSide Properties. For more information, see "Configuring Server-Side Properties" onpage 17.

11.To save your settings and close the Hortonworks Spark ODBC Driver Configurationtool, clickOK.

Configuring AuthenticationSomeSpark servers are configured to require authentication for access. To connect to aSpark server, you must configure the Hortonworks ODBC Driver with SQL Connector forApache Spark to use the authentication mechanism that matches the access requirementsof the server and provides the necessary credentials.

For information about how to determine the type of authentication your Spark serverrequires, see "Authentication Options" on page 44.

ODBC applications that connect to Shark Server 2 or Spark Thrift Server using a DSN canpass in authentication credentials by defining them in the DSN. To configure authenticationfor a connection that uses a DSN, use the ODBC Data Source Administrator.

Normally, applications that are not Shark Server 2 or Spark Thrift Server aware and thatconnect using a DSN-less connection do not have a facility for passing authenticationcredentials to the Hortonworks ODBC Driver with SQL Connector for Apache Spark for aconnection. However, the Hortonworks Spark ODBC Driver Configuration tool enables youto configure authentication without using a DSN.

Important: Credentials defined in a DSN take precedence over credentialsconfigured using the driver configuration tool. Credentials configured using thedriver configuration tool apply for all connections that are made using a DSN-lessconnection unless the client application is Shark Server 2 or Spark Thrift Serveraware and requests credentials from the user.

Using No Authentication

When connecting to a Spark server of type Shark Server, you must use No Authentication.When you use No Authentication, Binary is the only Thrift transport protocol that issupported.

Architecting the Future of Big Data

Hortonworks Inc. Page 11

To configure a connection without authentication:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In the Mechanism list, select No Authentication.3. If the Spark server is configured to use SSL, then click SSL Options to configure

SSL for the connection. For more information, see "Configuring SSL Verification" onpage 18.

4. To save your settings and close the dialog box, clickOK

Example connection string for Shark Server:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=1;AuthMech=0;Schema=Spark_database

Example connection string for Shark Server 2:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=0;Schema=Spark_database

Example connection string for Spark Thrift Server:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=3;AuthMech=0;Schema=Spark_database

Using Kerberos

Kerberos must be installed and configured before you can use this authenticationmechanism. For more information, see "Configuring Kerberos Authentication for Windows"on page 21.

Note: This authentication mechanism is available only for Shark Server 2 or SparkThrift Server on non-HDInsight distributions.

To configure Kerberos authentication:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In theMechanism list, select Kerberos.3. Choose one:

Architecting the Future of Big Data

Hortonworks Inc. Page 12

l To use the default realm defined in your Kerberos setup, leave the Realm fieldempty.

l Or, if your Kerberos setup does not define a default realm or if the realm of yourShark Server 2 or Spark Thrift Server host is not the default, then, in theRealmfield, type the Kerberos realm of the Shark Server 2 or Spark Thrift Server.

4. In theHost FQDN field, type the fully qualified domain name of the Shark Server 2 orSpark Thrift Server host.Note: To use the Spark server host name as the fully qualified domain name forKerberos authentication, in the Host FQDN  field, type _HOST.

5. In the Service Name field, type the service name of the Spark server.6. In the Thrift Transport list, select the transport protocol to use in the Thrift layer.

Important: When using this authentication mechanism, the Binary transportprotocol is not supported.

7. If the Spark server is configured to use SSL, then click SSL Options to configureSSL for the connection. For more information, see "Configuring SSL Verification" onpage 18.

8. To save your settings and close the dialog box, clickOK.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=1;Schema=Spark_database;KrbRealm=Kerberos_realm;KrbHostFQDN=domain_name;KrbServiceName=service_name

Using User Name

This authentication mechanism requires a user name but not a password. The user namelabels the session, facilitating database tracking.

Note: This authentication mechanism is available only for Shark Server 2 or SparkThrift Server on non-HDInsight distributions. Most default configurations of SharkServer 2 or Spark Thrift Server require User Name authentication.

To configure User Name authentication:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In theMechanism list, select User Name3. In theUser Name field, type an appropriate user name for accessing the Spark

server.

Architecting the Future of Big Data

Hortonworks Inc. Page 13

4. To save your settings and close the dialog box, clickOK.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=2;Schema=Spark_database;UID=user_name

Using User Name and Password

This authentication mechanism requires a user name and a password.

To configure User Name and Password authentication:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In theMechanism list, select User Name and Password3. In theUser Name field, type an appropriate user name for accessing the Spark

server.4. In the Password field, type the password corresponding to the user name you typed

in step 3.5. To save the password, select the Save Password (Encrypted) check box.

Important: The password is obscured, that is, not saved in plain text. However, it isstill possible for the encrypted password to be copied and used.

6. In the Thrift Transport list, select the transport protocol to use in the Thrift layer.7. If the Spark server is configured to use SSL, then click SSL Options to configure

SSL for the connection. For more information, see "Configuring SSL Verification" onpage 18.

8. To save your settings and close the dialog box, clickOK.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=3;Schema=Spark_database;UID=user_name;PWD=password

Using Windows Azure HDInsight Emulator

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Serverinstances running on Windows Azure HDInsight Emulator.

Architecting the Future of Big Data

Hortonworks Inc. Page 14

To configure a connection to a Spark server on Windows Azure HDInsightEmulator:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In theMechanism list, selectWindows Azure HDInsight Emulator.3. In theHTTP Path field, type the partial URL corresponding to the Spark server.4. In theUser Name field, type an appropriate user name for accessing the Spark

server.5. In the Password field, type the password corresponding to the user name you typed

in step 4.6. To save your settings and close the dialog box, clickOK.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=5;Schema=Spark_database;UID=user_name;PWD=password;HTTPPath=Spark_HTTP_path

Using Windows Azure HDInsight Service

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Server onHDInsight distributions.

To configure a connection to a Spark server on Windows Azure HDInsight Service:1. Choose one:

l To access authentication options for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, and then clickConfigure.

l Or, to access authentication options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool.

2. In theMechanism list, selectWindows Azure HDInsight Service.3. In theHTTP Path field, type the partial URL corresponding to the Spark server.4. In theUser Name field, type an appropriate user name for accessing the Spark

server.5. In the Password field, type the password corresponding to the user name you typed

in step 4.6. ClickHTTP Options, and in theHTTP Path field, type the partial URL corresponding

to the Spark server. Click OK to save your HTTP settings and close the dialog box.

Architecting the Future of Big Data

Hortonworks Inc. Page 15

Note: If necessary, you can create customHTTP headers. For more information,see "Configuring HTTP Options" on page 17.

7. Choose one:l To configure the driver to load SSL certificates from a specific PEM file, clickAdvanced Options and type the path to the file in the Trusted Certificatesfield.

l Or, to use the trusted CA certificates PEM file that is installed with the driver,leave the Trusted Certificates field empty.

8. Click SSL Options and configure SSL settings as needed. For more information, see"Configuring SSL Verification" on page 18.

9. ClickOK to save your SSL configuration and close the dialog box, and then clickOK tosave your authentication settings and close the dialog box.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=6;Schema=Spark_database;UID=user_name;PWD=password;HTTPPath=Spark_HTTP_path

Configuring Advanced OptionsYou can configure advanced options to modify the behavior of the driver.

To configure advanced options:1. Choose one:

l To access advanced options for a DSN, open the ODBC Data Source Admin-istrator where you created the DSN, then select the DSN, then clickConfigure,and then clickAdvanced Options.

l Or, to access advanced options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool, and then clickAdvancedOptions.

2. To disable the SQL Connector feature, select the Use Native Query check box.3. To defer query execution to SQLExecute, select the Fast SQLPrepare check box.4. To allow driver-wide configurations to take precedence over connection and DSN

settings, select theDriver Config Take Precedence check box.5. To use the asynchronous version of the API call against Spark for executing a query,

select the Use Async Exec check box.6. To retrieve the names of tables in a database by using the SHOW TABLES query,

select the Get Tables With Query check box.Note: This option is applicable onlywhen connecting to Shark Server 2 or SparkThrift Server.

Architecting the Future of Big Data

Hortonworks Inc. Page 16

7. To enable the driver to return SQL_WVARCHAR instead of SQL_VARCHAR forSTRINGand VARCHAR columns, and SQL_WCHAR instead of SQL_CHAR forCHAR columns, select the Unicode SQL character types check box.

8. To enable the driver to return the spark_system table for catalog function calls such asSQLTables and SQLColumns, select the Show System Table check box.

9. To handle Kerberos authentication using the SSPI plugin instead of MIT Kerberos bydefault, select theUse Only SSPI check box.

10.To enable the driver to automatically open a new session when the existing session isno longer valid, select the Invalid Session Auto Recover check box.Note: This option is applicable onlywhen connecting to Shark Server 2 or SparkThrift Server.

11. In theRows Fetched Per Block field, type the number of rows to be fetched perblock.

12. In theDefault String Column Length field, type the maximum data length forSTRING columns.

13. In theBinary Column Length field, type the maximum data length for BINARYcolumns.

14. In theDecimal Column Scale field, type the maximum number of digits to the right ofthe decimal point for numeric data types.

15. In theAsync Exec Poll Interval (ms) field, type the time in milliseconds betweeneach poll for the query execution status.

Note: This option is applicable only to HDInsight clusters.

16.To allow the common name of a CA-issued SSL certificate to not match the host nameof the Spark server, select the Allow Common Name Host Name Mismatch checkbox.Note: This option is applicable only to the User Name and Password (SSL) andHTTPS authentication mechanisms.

17.Choose one:l To configure the driver to load SSL certificates from a specific PEM file, type thepath to the file in the Trusted Certificates field.

l Or, to use the trusted CA certificates PEM file that is installed with the driver,leave the Trusted Certificates field empty.

Note: This option is applicable only to the User Name and Password (SSL),Windows Azure HDInsight Service, and HTTPS authentication mechanisms.

18. In the Socket Timeout field, type the number of seconds that an operation canremain idle before it is closed.Note: This option is applicable onlywhen asynchronous query execution is beingused against Shark Server 2 or Spark Thrift Server instances.

19.To save your settings and close the Advanced Options dialog box, clickOK.

Architecting the Future of Big Data

Hortonworks Inc. Page 17

Configuring Server-Side PropertiesYou can use the driver to apply configuration properties to the Spark server.

To configure server-side properties:1. Choose one:

l To configure server-side properties for a DSN, open the ODBC Data SourceAdministrator where you created the DSN, then select the DSN, then clickCon-figure, then click Advanced Options, and then click Server Side Properties.

l Or, to configure server-side properties for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool, then clickAdvancedOptions, and then click Server Side Properties.

2. To create a server-side property, click Add, then type appropriate values in the Keyand Value fields, and then clickOK.Note: For a list of all Hadoop and Spark server-side properties that yourimplementation supports, type set -v at the Spark CLI command line. You can alsoexecute the set -v query after connecting using the driver.

3. To edit a server-side property, select the property from the list, then click Edit, thenupdate the Key and Value fields as needed, and then clickOK.

4. To delete a server-side property, select the property from the list, and then clickRemove. In the confirmation dialog box, click Yes.

5. To configure the driver to apply each server-side property by executing a query whenopening a session to the Spark server, select theApply Server Side Propertieswith Queries check box.Or, to configure the driver to use amore efficient method for applying server-sideproperties that does not involve additional network round-tripping, clear the ApplyServer Side Properties with Queries check box.

Note: The more efficient method is not available for Shark Server, and it might notbe compatible with some Shark Server 2 or Spark Thrift Server builds. If the server-side properties do not take effect when the check box is clear, then select the checkbox.

6. To force the driver to convert server-side property key names to all lower casecharacters, select theConvert Key Name to Lower Case check box.

7. To save your settings and close the Server Side Properties dialog box, clickOK.

Configuring HTTP OptionsYou can configure options such as custom headers when using the HTTP transport protocolin the Thrift layer.

Architecting the Future of Big Data

Hortonworks Inc. Page 18

To configure HTTP options:1. Chose one:

l If you are configuring HTTP for a DSN, open the ODBC Data Source Admin-istrator where you created the DSN, then select the DSN, then clickConfigure,and then ensure that the Thrift Transport option is set to HTTP.

l Or, if you are configuring HTTP for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool and then ensure that the ThriftTransport option is set to HTTP.

2. To access HTTP options, clickHTTP Options.Note: The HTTP options are available only when the Thrift Transport option is setto HTTP.

3. In theHTTP Path field, type the partial URL corresponding to the Spark server.4. To create a custom HTTP header, clickAdd, then type appropriate values in theKey

and Value fields, and then clickOK.5. To edit a custom HTTP header, select the header from the list, then click Edit, then

update the Key and Value fields as needed, and then clickOK.6. To delete a custom HTTP header, select the header from the list, and then click

Remove. In the confirmation dialog box, click Yes.7. To save your settings and close the HTTP Options dialog box, click OK.

Configuring SSL VerificationYou can configure verification between the client and the Spark server over SSL.

To configure SSL verification:1. Choose one:

l To access SSL options for a DSN, open the ODBC Data Source Administratorwhere you created the DSN, then select the DSN, then click Configure, andthen click SSL Options.

l Or, to access advanced options for a DSN-less connection, open the Hor-tonworks Spark ODBC Driver Configuration tool, and then click SSL Options.

2. Select the Enable SSL check box.3. To allow self-signed certificates from the server, select the Allow Self-signed Server

Certificate check box.4. To allow the common name of a CA-issued SSL certificate to not match the host name

of the Spark server, select the Allow Common Name Host Name Mismatch checkbox.

5. Choose one:l To configure the driver to load SSL certificates from a specific PEM file when

Architecting the Future of Big Data

Hortonworks Inc. Page 19

verifying the server, specify the full path to the file in the Trusted Certificatesfield.

l Or, to use the trusted CA certificates PEM file that is installed with the driver,leave the Trusted Certificates field empty.

6. To configure two-way SSL verification, select the Two Way SSL check box and thendo the following:

a) In theClient Certificate File field, specify the full path of the PEM filecontaining the client's certificate.

b) In theClient Private Key File field, specify the full path of the file containing theclient's private key.

c) If the private key file is protected with a password, type the password in theClient Private Key Password field. To save the password, select the SavePassword (Encrypted) check box.Important: The password is obscured, that is, not saved in plain text.However, it is still possible for the encrypted password to be copied andused.

7. To save your settings and close the SSL Options dialog box, clickOK.

Configuring Logging OptionsTo help troubleshoot issues, you can enable logging. In addition to functionality provided inthe Hortonworks ODBC Driver with SQL Connector for Apache Spark, the ODBC DataSource Administrator provides tracing functionality.

Important: Only enable logging or tracing long enough to capture an issue.Logging or tracing decreases performance and can consume a large quantity of diskspace.

The driver allows you to set the amount of detail included in log files. Table 1 lists the logginglevels provided by the Hortonworks ODBC Driver with SQL Connector for Apache Spark, inorder from least verbose to most verbose.

Logging Level Description

OFF Disables all logging.

FATAL Logs very severe error events that will lead the driver to abort.

ERROR Logs error events that might still allow the driver to continuerunning.

Table 1. Hortonworks ODBC Driver with SQL Connector for Apache Spark Logging Levels

Architecting the Future of Big Data

Hortonworks Inc. Page 20

Logging Level Description

WARNING Logs potentially harmful situations.

INFO Logs general information that describes the progress of thedriver.

DEBUG Logs detailed information that is useful for debugging the driver.

TRACE Logsmore detailed information than the DEBUG level.

To enable driver logging:1. To access logging options, open the ODBC Data Source Administrator where you

created the DSN, then select the DSN, then clickConfigure, and then click LoggingOptions.

2. From the Log Level drop-down list, select the desired level of information to include inlog files.

3. In the Log Path field, specify the full path to the folder where you want to save logfiles.

4. In theMax Number Files field, type the maximum number of log files to keep.Note: After the maximum number of log files is reached, each time an additional fileis created, the driver deletes the oldest log file.

5. In theMax File Size field, type the maximum size of each log file in megabytes (MB).Note: After the maximum file size is reached, the driver creates a new file andcontinues logging.

6. ClickOK.7. Restart your ODBC application to ensure that the new settings take effect.

The Hortonworks ODBC Driver with SQL Connector for Apache Spark produces a log filenamed SparkODBC_driver.log at the location that you specify in the Log Path field.

To disable driver logging:1. Open the ODBC Data Source Administrator where you created the DSN, then select

the DSN, then clickConfigure, and then click Logging Options.2. From the Log Level drop-down list, select LOG_OFF.3. ClickOK.

To start tracing using the ODBC Data Source Administrator:1. In the ODBC Data Source Administrator, click the Tracing tab.

Architecting the Future of Big Data

Hortonworks Inc. Page 21

2. In the Log File Path area, clickBrowse. In the Select ODBC Log File dialog box,browse to the location where you want to save the log file, then type a descriptive filename in the File name field, and then click Save.

3. On the Tracing tab, click Start Tracing Now.

To stop ODBC Data Source Administrator tracing:On the Tracing tab in the ODBC Data Source Administrator, click Stop Tracing Now.

For more information about tracing using the ODBC Data Source Administrator, see thearticle How to Generate an ODBC Trace with ODBC Data Source Administrator athttp://support.microsoft.com/kb/274551.

Configuring Kerberos Authentication for WindowsActive Directory

The Hortonworks ODBC Driver with SQL Connector for Apache Spark supports ActiveDirectory Kerberos on Windows. There are two prerequisites for using Active DirectoryKerberos on Windows:

l MIT Kerberos is not installed on the client Windows machine.l The MIT Kerberos Hadoop realm has been configured to trust the Active Directoryrealm so that users in the Active Directory realm can access services in the MIT Ker-beros Hadoop realm. For more information, see Setting up One-Way Trust with Act-ive Directory in the Hortonworks documentation athttp://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.1.7/bk_installing_manu-ally_book/content/ch23s05.html.

MIT Kerberos

Downloading and InstallingMIT Kerberos for Windows 4.0.1

For information about Kerberos and download links for the installer, see the MIT Kerberoswebsite at http://web.mit.edu/kerberos/.

To download and install MIT Kerberos for Windows 4.0.1:1. Download the appropriate Kerberos installer:

l For a 64-bit computer, use the following download link from the MIT Kerberoswebsite: http://web.mit.edu/kerberos/dist/kfw/4.0/kfw-4.0.1-amd64.msi.

l For a 32-bit computer, use the following download link from the MIT Kerberoswebsite: http://web.mit.edu/kerberos/dist/kfw/4.0/kfw-4.0.1-i386.msi.

Note: The 64-bit installer includes both 32-bit and 64-bit libraries. The 32-bitinstaller includes 32-bit libraries only.

2. To run the installer, double-click the .msi file that you downloaded in step 1.

Architecting the Future of Big Data

Hortonworks Inc. Page 22

3. Follow the instructions in the installer to complete the installation process.4. When the installation completes, click Finish.

Setting Up the KerberosConfiguration File

Settings for Kerberos are specified through a configuration file. You can set up theconfiguration file as a .INI file in the default location, which is theC:\ProgramData\MIT\Kerberos5 directory, or as a .CONF file in a custom location.

Normally, the C:\ProgramData\MIT\Kerberos5 directory is hidden. For information aboutviewing and using this hidden directory, refer to Microsoft Windows documentation.

Note: For more information on configuring Kerberos, refer to the MIT Kerberosdocumentation.

To set up the Kerberos configuration file in the default location:1. Obtain a krb5.conf configuration file . You can obtain this file from your Kerberos

administrator, or from the /etc/krb5.conf folder on the computer that is hosting theShark Server 2 or Spark Thrift Server instance.

2. Rename the configuration file from krb5.conf to krb5.ini.3. Copy the krb5.ini file to the C:\ProgramData\MIT\Kerberos5 directory and overwrite

the empty sample file.

To set up the Kerberos configuration file in a custom location:1. Obtain a krb5.conf configuration file . You can obtain this file from your Kerberos

administrator, or from the /etc/krb5.conf folder on the computer that is hosting theShark Server 2 or Spark Thrift Server instance.

2. Place the krb5.conf file in an accessible directory and make note of the full path name.3. Open your computer Properties window:

l If you are using Windows 7 or earlier, click the Start button , then right-clickComputer, and then click Properties.

l Or, if you are using Windows 8 or later, right-click This PC on the Start screen,and then click Properties.

4. ClickAdvanced System Settings.5. In the System Properties dialog box, click theAdvanced tab and then click

Environment Variables.6. In the Environment Variables dialog box, under the System variables list, click New.7. In the New SystemVariable dialog box, in the Variable Name field, type KRB5_

CONFIG.8. In the Variable Value field, type the absolute path to the krb5.conf file from step 2.9. ClickOK to save the new variable.10.Make sure that the variable is listed in the System variables list..

Architecting the Future of Big Data

Hortonworks Inc. Page 23

11.ClickOK to close the Environment Variables dialog box, and then clickOK to close theSystemProperties dialog box.

Setting Up the KerberosCredential Cache File

Kerberos uses a credential cache to store and manage credentials.

To set up the Kerberos credential cache file:1. Create a directory where you want to save the Kerberos credential cache file. For

example, create a directory named C:\temp.2. Open your computer Properties window:

l If you are using Windows 7 or earlier, click the Start button , then right-clickComputer, and then click Properties.

l Or, if you are using Windows 8 or later, right-click This PC on the Start screen,and then click Properties.

3. ClickAdvanced System Settings4. In the System Properties dialog box, click theAdvanced tab and then click

Environment Variables.5. In the Environment Variables dialog box, under the SystemVariables list, click New.6. In the New SystemVariable dialog box, in the Variable Name field, type

KRB5CCNAME.7. In the Variable Value field, type the path to the folder you created in step 1, and then

append the file name krb5cache. For example, if you created the folder C:\temp instep 1, then type C:\temp\krb5cache.Note: krb5cache is a file (not a directory) that is managed by the Kerberossoftware, and it should not be created by the user. If you receive a permission errorwhen you first use Kerberos, ensure that the krb5cache file does not already existas a file or a directory.

8. ClickOK to save the new variable.9. Make sure that the variable appears in the SystemVariables list.10.ClickOK to close the Environment Variables dialog box, and then clickOK to close the

SystemProperties dialog box.11.To ensure that Kerberos uses the new settings, restart your computer.

Obtaining a Ticket for a Kerberos Principal

A principal refers to a user or service that can authenticate to Kerberos. To authenticate toKerberos, a principal must obtain a ticket by using a password or a keytab file. You canspecify a keytab file to use, or use the default keytab file of your Kerberos configuration.

To obtain a ticket for a Kerberos principal using a password:1. OpenMIT Kerberos Ticket Manager.

Architecting the Future of Big Data

Hortonworks Inc. Page 24

2. In the MIT Kerberos Ticket Manager, clickGet Ticket.3. In the Get Ticket dialog box, type your principal name and password, and then click

OK.

If the authentication succeeds, then your ticket information appears in the MIT KerberosTicket Manager.

To obtain a ticket for a Kerberos principal using a keytab file:1. Open a command prompt:

l If you are using Windows 7 or earlier, click the Start button , then click AllPrograms, then clickAccessories, and then clickCommand Prompt.

l If you are using Windows 8 or later, click the arrow button at the bottom of theStart screen, then find theWindows System program group, and then clickCom-mand Prompt .

2. In the Command Prompt, type a command using the following syntax:kinit -k -t keytab_path principal

keytab_path is the full path to the keytab file. For example:C:\mykeytabs\myUser.keytabprincipal is the Kerberos user principal to use for authentication. For example:[email protected]

3. If the cache location KRB5CCNAME is not set or used, then use the -c option of thekinit command to specify the location of the credential cache. In the command, the -cargument must appear last. For example:kinit -k -t C:\mykeytabs\myUser.keytab [email protected] -cC:\ProgramData\MIT\krbcache

Krbcache is the Kerberos cache file, not a directory.

To obtain a ticket for a Kerberos principal using the default keytab file:

Note: For information about configuring a default keytab file for your Kerberosconfiguration, refer to the MIT Kerberos documentation.1. Open a command prompt:

l If you are using Windows 7 or earlier, click the Start button , then click AllPrograms, then clickAccessories, and then clickCommand Prompt.

l If you are using Windows 8 or later, click the arrow button at the bottom of theStart screen, then find theWindows System program group, and then clickCom-mand Prompt .

2. In the Command Prompt, type a command using the following syntax:kinit -k principal

principal is the Kerberos user principal to use for authentication, for example,[email protected].

Architecting the Future of Big Data

Hortonworks Inc. Page 25

3. If the cache location KRB5CCNAME is not set or used, then use the -c option of thekinit command to specify the location of the credential cache. In the command, the -cargument must appear last. For example:kinit -k -t C:\mykeytabs\myUser.keytab [email protected] -cC:\ProgramData\MIT\krbcache

Krbcache is the Kerberos cache file, not a directory.

Verifying the Version NumberIf you need to verify the version of the Hortonworks ODBC Driver with SQL Connector forApache Spark that is installed on your Windows machine, you can find the version number inthe ODBC Data Source Administrator.

To verify the version number:1. Open the ODBC Administrator:

l If you are using Windows 7 or earlier, click the Start button , then click AllPrograms, then click the Hortonworks Spark ODBC Driver 1.1 programgroup corresponding to the bitness of the client application accessing data inSpark, and then clickODBC Administrator

l Or, if you are using Windows 8 or later, on the Start screen, type ODBC admin-istrator, and then click theODBC Administrator search result correspondingto the bitness of the client application accessing data in Spark.

2. Click theDrivers tab and then find the Hortonworks Spark ODBC Driver in the list ofODBC drivers that are installed on your system. The version number is displayed inthe Version column.

Architecting the Future of Big Data

Hortonworks Inc. Page 26

Linux DriverLinux System RequirementsYou install the Hortonworks ODBC Driver with SQL Connector for Apache Spark on clientmachines that access data stored in a Hadoop cluster with the Spark service installed andrunning. Each machine that you install the driver on must meet the followingminimumsystem requirements:

l One of the following distributions:o Red Hat® Enterprise Linux® (RHEL) 5 or 6o CentOS 5, 6, or 7o SUSE Linux Enterprise Server (SLES) 11 or 12

l 150 MB of available disk spacel One of the followingODBC driver managers installed:

o iODBC 3.52.7 or latero unixODBC 2.2.14 or later

The driver supports Apache Spark versions 0.8 through 1.5.

Installing the DriverThere are two versions of the driver for Linux:

l spark-odbc-native-32bit-Version-Release.LinuxDistro.i686.rpm for the 32-bitdriver

l spark-odbc-native-Version-Release.LinuxDistro.x86_64.rpm for the 64-bit driver

Version is the version number of the driver, and Release is the release number for thisversion of the driver. LinuxDistro is either el5 or el6. For SUSE, the LinuxDistro placeholderis empty.

The bitness of the driver that you select should match the bitness of the client applicationaccessing your Hadoop-based data. For example, if the client application is 64-bit, then youshould install the 64-bit driver. Note that 64-bit editions of Linux support both 32- and 64-bitapplications. Verify the bitness of your intended application and install the appropriateversion of the driver.

Important: Ensure that you install the driver using the RPM corresponding to yourLinux distribution.

Architecting the Future of Big Data

Hortonworks Inc. Page 27

The Hortonworks ODBC Driver with SQL Connector for Apache Spark driver files areinstalled in the following directories:

l /usr/lib/spark/lib/native/sparkodbc contains release notes, the Hortonworks ODBCDriver with SQL Connector for Apache Spark User Guide in PDF format, and aReadme.txt file that provides plain text installation and configuration instructions.

l /usr/lib/spark/lib/native/sparkodbc/ErrorMessages contains error message filesrequired by the driver.

l /usr/lib/spark/lib/native/Linux-i386-32 contains the 32-bit driver and thehortonworks.sparkodbc.ini configuration file.

l /usr/lib/spark/lib/native/Linux-amd64-64 contains the 64-bit driver and thehortonworks.sparkodbc.ini configuration file.

To install the Hortonworks ODBC Driver with SQL Connector for Apache Spark:1. Choose one:

l In Red Hat Enterprise Linux or CentOS, log in as the root user, then navigate tothe folder containing the driver RPM packages to install, and then type the fol-lowing at the command line, where RPMFileName is the file name of the RPMpackage containing the version of the driver that you want to install:yum --nogpgcheck localinstall RPMFileName

l In SUSE Linux Enterprise Server, log in as the root user, then navigate to thefolder containing the driver RPM packages to install, and then type the followingat the command line, where RPMFileName is the file name of the RPM packagecontaining the version of the driver that you want to install:zypper install RPMFileName

The Hortonworks ODBC Driver with SQL Connector for Apache Spark depends on thefollowing resources:

l cyrus-sasl-2.1.22-7 or abovel cyrus-sasl-gssapi-2.1.22-7 or abovel cyrus-sasl-plain-2.1.22-7 or above

If the package manager in your Linux distribution cannot resolve the dependenciesautomatically when installing the driver, then download and manually install the packagesrequired by the version of the driver that you want to install.

Setting the LD_LIBRARY_PATH Environment VariableThe LD_LIBRARY_PATH environment variable must include the paths to the installedODBC driver manager libraries.

For example, if ODBC driver manager libraries are installed in /usr/local/lib, then set LD_LIBRARY_PATH as follows:export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib

Architecting the Future of Big Data

Hortonworks Inc. Page 28

For information about how to set environment variables permanently, refer to your Linuxshell documentation.

For information about creating ODBC connections using the Hortonworks ODBC Driver withSQL Connector for Apache Spark, see "Configuring ODBC Connections for Non-WindowsPlatforms" on page 31.

Verifying the Version NumberIf you need to verify the version of the Hortonworks ODBC Driver with SQL Connector forApache Spark that is installed on your Linux machine, you can query the version numberthrough the command-line interface.

To verify the version number:Depending on your version of Linux, at the command prompt, run one of the followingcommands:yum list | grep spark-odbc-native

rpm -qa | grep spark-odbc-native

The command returns information about the Hortonworks ODBC Driver with SQLConnector for Apache Spark that is installed on your machine, including the version number.

Architecting the Future of Big Data

Hortonworks Inc. Page 29

Mac OS X DriverMac OS X System RequirementsYou install the Hortonworks ODBC Driver with SQL Connector for Apache Spark on clientmachines that access data stored in a Hadoop cluster with the Spark service installed andrunning. Each machine that you install the driver on must meet the followingminimumsystem requirements:

l MacOS X version 10.9 or 10.10l 100 MB of available disk spacel iODBC 3.52.7 or later

The driver supports both 32- and 64-bit client applications.

The driver supports Apache Spark versions 0.8 through 1.5.

Installing the DriverThe Hortonworks ODBC Driver with SQL Connector for Apache Spark driver files areinstalled in the following directories:

l /opt/hortonworks/sparkodbc contains release notes and theHortonworks ODBCDriver with SQL Connector for Apache Spark User Guide in PDF format.

l /opt/hortonworks/sparkodbc/ErrorMessages contains error message filesrequired by the driver.

l /opt/hortonworks/sparkodbc/Setup contains sample configuration files namedodbc.ini and odbcinst.ini.

l /opt/hortonworks/sparkodbc/lib/universal contains the driver binaries and thehortonworks.sparkodbc.ini configuration file.

To install the Hortonworks ODBC Driver with SQL Connector for Apache Spark:1. Double-click spark-odbc-native.dmg to mount the disk image.2. Double-click spark-odbc-native.pkg to run the installer.3. In the installer, click Continue.4. On the Software License Agreement screen, click Continue, and when the prompt

appears, clickAgree if you agree to the terms of the License Agreement.5. Optionally, to change the installation location, clickChange Install Location, then

select the desired location, and then clickContinue.6. To accept the installation location and begin the installation, click Install.7. When the installation completes, clickClose.

Architecting the Future of Big Data

Hortonworks Inc. Page 30

Setting the DYLD_LIBRARY_PATH Environment VariableThe DYLD_LIBRARY_PATH environment variable must include the paths to the installedODBC driver manager libraries.

For example, if ODBC driver manager libraries are installed in /usr/local/lib, then setDYLD_LIBRARY_PATH as follows:export DYLD_LIBRARY_PATH=$DYLD_LIBRARY_PATH:/usr/local/lib

For information about how to set environment variables permanently, refer to your Mac OSX shell documentation.

For information about creating ODBC connections using the Hortonworks ODBC Driver withSQL Connector for Apache Spark, see "Configuring ODBC Connections for Non-WindowsPlatforms" on page 31.

Verifying the Version NumberIf you need to verify the version of the Hortonworks ODBC Driver with SQL Connector forApache Spark that is installed on your Mac OS X machine, you can query the versionnumber through the Terminal.

To verify the version number:At the Terminal, run the following command:pkgutil --info com.hortonworks.sparkodbc

The command returns information about the Hortonworks ODBC Driver with SQLConnector for Apache Spark that is installed on your machine, including the version number.

Architecting the Future of Big Data

Hortonworks Inc. Page 31

Configuring ODBC Connections for Non-Win-dows PlatformsThe following sections describe how to configure ODBC connections when using theHortonworks ODBC Driver with SQL Connector for Apache Spark with non-Windowsplatforms:

l "Configuration Files" on page 31l "Sample Configuration Files" on page 32l "Configuring the Environment" on page 32l "Defining DSNs in odbc.ini" on page 33l "Specifying ODBC drivers in odbcinst.ini" on page 34l "Configuring Driver Settings in hortonworks.sparkodbc.ini" on page 35l "Configuring Authentication" on page 36l "Configuring SSL Verification" on page 39l "Configuring Logging Options" on page 42

Configuration FilesODBC driver managers use configuration files to define and configure ODBC data sourcesand drivers. By default, the following configuration files residing in the user’s home directoryare used:

l .odbc.ini is used to define ODBC data sources, and it is required for DSNs.l .odbcinst.ini is used to define ODBC drivers, and it is optional.

Also, by default the Hortonworks ODBC Driver with SQL Connector for Apache Spark isconfigured using the hortonworks.sparkodbc.ini file, which is located in one of thefollowing directories depending on the version of the driver that you are using:

l /usr/lib/spark/lib/native/Linux-i386-32 for the 32-bit driver on Linuxl /usr/lib/spark/lib/native/Linux-amd64-64 for the 64-bit driver on Linuxl /opt/hortonworks/sparkodbc/lib/universal for the driver on MacOS X

The hortonworks.sparkodbc.ini file is required.

Note: The hortonworks.sparkodbc.ini file in the /lib subfolder providesdefault settings for most configuration options available in the Hortonworks ODBCDriver with SQL Connector for Apache Spark.

You can set driver configuration options in your odbc.ini and hortonworks.sparkodbc.ini files.Configuration options set in a hortonworks.sparkodbc.ini file apply to all connections,whereas configuration options set in an odbc.ini file are specific to a connection.Configuration options set in odbc.ini take precedence over configuration options set in

Architecting the Future of Big Data

Hortonworks Inc. Page 32

hortonworks.sparkodbc.ini. For information about the configuration options available forcontrolling the behavior of DSNs that are using the Hortonworks ODBC Driver with SQLConnector for Apache Spark, see "Driver Configuration Options" on page 55.

Sample Configuration FilesThe driver installation contains the following sample configuration files in the Setupdirectory:

l odbc.inil odbcinst.ini

These sample configuration files provide preset values for settings related to theHortonworks ODBC Driver with SQL Connector for Apache Spark.

The names of the sample configuration files do not begin with a period (.) so that theyappear in directory listings by default. A file name beginning with a period (.) is hidden. Forodbc.ini and odbcinst.ini, if the default location is used, then the file names mustbegin with a period (.).

If the configuration files do not exist in the home directory, then you can copy the sampleconfiguration files to the home directory, and then rename the files. If the configuration filesalready exist in the home directory, then use the sample configuration files as a guide tomodify the existing configuration files.

Configuring the EnvironmentOptionally, you can use three environment variables, ODBCINI, ODBCSYSINI, andHORTONWORKSSPARKINI, to specify different locations for the odbc.ini,odbcinst.ini, and hortonworks.sparkodbc.ini configuration files by doing thefollowing:

l Set ODBCINI to point to your odbc.ini file.l Set ODBCSYSINI to point to the directory containing the odbcinst.ini file.l Set HORTONWORKSSPARKINI to point to your hortonworks.sparkodbc.inifile.

For example, if your odbc.ini and hortonworks.sparkodbc.ini files are located in/etc and your odbcinst.ini file is located in /usr/local/odbc, then set the environmentvariables as follows:export ODBCINI=/etc/odbc.ini

export ODBCSYSINI=/usr/local/odbc

export HORTONWORKSSPARKINI=/etc/hortonworks.sparkodbc.ini

The following search order is used to locate the hortonworks.sparkodbc.ini file:

Architecting the Future of Big Data

Hortonworks Inc. Page 33

1. If the HORTONWORKSSPARKINI environment variable is defined, then the driversearches for the file specified by the environment variable.Important: HORTONWORKSSPARKINI must specify the full path, including thefile name.

2. The directory containing the driver’s binary is searched for a file namedhortonworks.sparkodbc.ini (not beginning with a period).

3. The current working directory of the application is searched for a file namedhortonworks.sparkodbc.ini (not beginning with a period).

4. The directory ~/ (that is, $HOME) is searched for a hidden file named.hortonworks.sparkodbc.ini.

5. The directory /etc is searched for a file named hortonworks.sparkodbc.ini(not beginning with a period).

Defining DSNs in odbc.iniODBC Data Source Names (DSNs) are defined in the odbc.ini configuration file. This fileis divided into several sections:

l [ODBC] is optional. This section is used to control global ODBC configuration, such asODBC tracing.

l [ODBC Data Sources] is required. This section lists the DSNs and associatesthemwith a driver.

l A section having the same name as the data source specified in the [ODBC DataSources] section is required to configure the data source.

The following is an example of an odbc.ini configuration file for Linux:[ODBC Data Sources]

Sample Hortonworks Spark ODBC DSN 32=Hortonworks Spark ODBC Driver32-bit

[Sample Hortonworks Spark DSN 32]

Driver=/usr/lib/spark/lib/native/Linux-i386-32/libhortonworkssparkodbc32.so

HOST=MySparkServer

PORT=10000

MySparkServer is the IP address or host name of the Spark server.

The following is an example of an odbc.ini configuration file for Mac OS X:[ODBC Data Sources]

Sample Hortonworks Spark DSN=Hortonworks Spark ODBC Driver

[Sample Hortonworks Spark DSN]

Architecting the Future of Big Data

Hortonworks Inc. Page 34

Driver=/opt/hortonworks/sparkodbc/lib/universal/libhortonworkssparkodbc.dylib

HOST=MySparkServer

PORT=10000

MySparkServer is the IP address or host name of the Spark server.

To create a Data Source Name:1. In a text editor, open the odbc.ini configuration file.2. In the [ODBC Data Sources] section, add a new entry by typing the Data Source

Name (DSN), then an equal sign (=), and then the driver name.3. Add a new section to the file, with a section name that matches the DSN you specified

in step 2, and then add configuration options to the section. Specify the configurationoptions as key-value pairs.Note: Shark Server does not support authentication. Most default configurations ofShark Server 2 or Spark Thrift Server require User Name authentication, which youconfigure by setting the AuthMech key to 2. To verify the authentication mechanismthat you need to use for your connection, check the configuration of your Hadoop /Spark distribution. For more information, see "Authentication Options" on page 44.

4. Save the odbc.ini configuration file.

For information about the configuration options available for controlling the behavior ofDSNs that are using the Hortonworks ODBC Driver with SQL Connector for Apache Spark,see "Driver Configuration Options" on page 55.

Specifying ODBC drivers in odbcinst.iniODBC drivers are defined in the odbcinst.ini configuration file. This configuration file isoptional because drivers can be specified directly in the odbc.ini configuration file, asdescribed in "Defining DSNs in odbc.ini" on page 33.

The odbcinst.ini file is divided into the following sections:l [ODBC Drivers] lists the names of all the installed ODBC drivers.l For each driver, a section having the same name as the driver name specified in the[ODBC Drivers] section lists the driver attributes and values.

The following is an example of an odbcinst.ini configuration file for Linux:[ODBC Drivers]

Hortonworks Spark ODBC Driver 32-bit=Installed

Hortonworks Spark ODBC Driver 64-bit=Installed

[Hortonworks Spark ODBC Driver 32-bit]

Description=Hortonworks Spark ODBC Driver (32-bit)

Architecting the Future of Big Data

Hortonworks Inc. Page 35

Driver=/usr/lib/spark/lib/native/Linux-i386-32/libhortonworkssparkodbc32.so

[Hortonworks Spark ODBC Driver 64-bit]

Description=Hortonworks Spark ODBC Driver (64-bit)

Driver=/usr/lib/spark/lib/native/Linux-amd64-64/libhortonworkssparkodbc64.so

The following is an example of an odbcinst.ini configuration file for Mac OS X:[ODBC Drivers]

Hortonworks Spark ODBC Driver=Installed

[Hortonworks Spark ODBC Driver]

Description=Hortonworks Spark ODBC Driver

Driver=/opt/hortonworks/sparkodbc/lib/universal/libhortonworkssparkodbc.dylib

To define a driver:1. In a text editor, open the odbcinst.ini configuration file.2. In the [ODBC Drivers] section, add a new entry by typing the driver name and then

typing =Installed.Note: Give the driver a symbolic name that you want to use to refer to the driver inconnection strings or DSNs.

3. Add a new section to the file with a name that matches the driver name you typed instep 2, and then add configuration options to the section based on the sampleodbcinst.ini file provided in the Setup directory. Specify the configuration optionsas key-value pairs.

4. Save the odbcinst.ini configuration file.

Configuring Driver Settings in hortonworks.sparkodbc.iniThe hortonworks.sparkodbc.ini file contains configuration settings for the HortonworksODBC Driver with SQL Connector for Apache Spark. Settings that you define in thehortonworks.sparkodbc.ini file apply to all connections that use the driver.

You do not need to modify the settings in the hortonworks.sparkodbc.ini file to usethe driver and connect to your data source.

However, to help troubleshoot issues, you can configure thehortonworks.sparkodbc.ini file to enable logging in the driver. For information aboutconfiguring logging, see "Configuring Logging Options" on page 42.

Architecting the Future of Big Data

Hortonworks Inc. Page 36

Configuring AuthenticationSomeSpark servers are configured to require authentication for access. To connect to aSpark server, you must configure the Hortonworks ODBC Driver with SQL Connector forApache Spark to use the authentication mechanism that matches the access requirementsof the server and provides the necessary credentials.

For information about how to determine the type of authentication your Spark serverrequires, see "Authentication Options" on page 44.

You can select the type of authentication to use for a connection by defining the AuthMechconnection attribute in a connection string or in a DSN (in the odbc.ini file). Depending onthe authentication mechanism you use, there may be additional connection attributes thatyou must define. For more information about the attributes involved in configuringauthentication, see "Driver Configuration Options" on page 55.

Using No Authentication

When connecting to a Spark server of type Shark Server, you must use No Authentication.When you use No Authentication, Binary is the only Thrift transport protocol that issupported.

To configure a connection without authentication:1. Set the AuthMech connection attribute to 0.2. If the Spark server is configured to use SSL, then configure SSL for the connection.

For more information, see "Configuring SSL Verification" on page 39.

Example connection string for Shark Server:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=1;AuthMech=0;Schema=Spark_database

Example connection string for Shark Server 2:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=0;Schema=Spark_database

Example connection string for Spark Thrift Server:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=3;AuthMech=0;Schema=Spark_database

Kerberos

Kerberos must be installed and configured before you can use this authenticationmechanism. For more information, refer to the MIT Kerberos documentation.

To configure Kerberos authentication:1. Set the AuthMech connection attribute to 1.

Architecting the Future of Big Data

Hortonworks Inc. Page 37

2. To use the default realm defined in your Kerberos setup, do not set the KrbRealmattribute.If your Kerberos setup does not define a default realm or if the realm of your Sparkserver is not the default, then set the appropriate realm using the KrbRealm attribute.

3. Set the KrbHostFQDN attribute to the fully qualified domain name of the Shark Server2 or Spark Thrift Server host.Note: To use the Spark server host name as the fully qualified domain name forKerberos authentication, set KrbHostFQDN to _HOST

4. Set the KrbServiceName attribute to the service name of the Shark Server 2 or SparkThrift Server.

5. Set the ThriftTransport connection attribute to the transport protocol to use in theThrift layer.Important: When using this authentication mechanism, Binary (ThriftTransport=0) is not supported.

6. If the Spark server is configured to use SSL, then configure SSL for the connection.For more information, see "Configuring SSL Verification" on page 39.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=1;Schema=Spark_database;KrbRealm=Kerberos_realm;KrbHostFQDN=domain_name;KrbServiceName=service_name

Using User Name

This authentication mechanism requires a user name but does not require a password. Theuser name labels the session, facilitating database tracking.

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Server onnon-HDInsight distributions. Most default configurations of require User Nameauthentication. When you use User Name authentication, SSL is not supported and SASL isthe only Thrift transport protocol available.

To configure User Name authentication:1. Set the AuthMech connection attribute to 2.2. Set the UID attribute to an appropriate user name for accessing the Spark server.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=2;Schema=Spark_database;UID=user_name

Architecting the Future of Big Data

Hortonworks Inc. Page 38

Using User Name and Password

This authentication mechanism requires a user name and a password.

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Server onnon-HDInsight distributions.

To configure User Name and Password authentication:1. Set the AuthMech connection attribute to 3.2. Set the UID attribute to an appropriate user name for accessing the Spark server.3. Set the PWD attribute to the password corresponding to the user name you provided

in step 2.4. Set the ThriftTransport connection attribute to the transport protocol to use in the

Thrift layer.5. If the Spark server is configured to use SSL, then configure SSL for the connection.

For more information, see "Configuring SSL Verification" on page 39.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=3;Schema=Spark_database;UID=user_name;PWD=password

Using Windows Azure HDInsight Emulator

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Serverinstances running on Windows Azure HDInsight Emulator. When you use this authenticationmechanism, SSL is not supported and HTTP is the only Thrift transport protocol available.

To configure a connection to a Spark server on Windows Azure HDInsightEmulator:1. Set the AuthMech connection attribute to 5.2. Set the HTTPPath attribute to the partial URL corresponding to the Spark server.3. Set the UID attribute to an appropriate user name for accessing the Spark server.4. Set the PWD attribute to the password corresponding to the user name you provided

in step 3.5. If necessary, you can create custom HTTP headers. For more information, see

"http.header." on page 72.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=5;Schema=Spark_database;UID=user_name;PWD=password;HTTPPath=Spark_HTTP_path

Architecting the Future of Big Data

Hortonworks Inc. Page 39

Using Windows Azure HDInsight Service

This authentication mechanism is available only for Shark Server 2 or Spark Thrift Server onHDInsight distributions. When you use this authentication mechanism, you must enableSSL and HTTP is the only Thrift transport protocol available.

To configure a connection to a Spark server on Windows Azure HDInsight Service:1. Set the AuthMech connection attribute to 6.2. Set the HTTPPath attribute to the partial URL corresponding to the Spark server.3. Set the UID attribute to an appropriate user name for accessing the Spark server.4. Set the PWD attribute to the password corresponding to the user name you typed in

step 3.5. If necessary, you can create custom HTTP headers. For more information, see

"http.header." on page 72.6. Configure SSL settings as needed. For more information, see "Configuring

SSL Verification" on page 39.7. Choose one:

l To configure the driver to load SSL certificates from a specific file, set the Trus-tedCerts attribute to the path of the file.

l Or, to use the trusted CA certificates PEM file that is installed with the driver, donot specify a value for the TrustedCerts attribute.

Example connection string:Driver=Hortonworks Spark ODBC Driver;Host=server_host;Port=10000;SparkServerType=2;AuthMech=6;Schema=Spark_database;UID=user_name;PWD=password;HTTPPath=Spark_HTTP_path

Configuring SSL VerificationYou can configure verification between the client and the Spark server over SSL.

To configure SSL verification:1. Open the odbc.ini configuration file in a text editor.2. To enable SSL, set the SSL attribute to 1.3. To allow self-signed certificates from the server, set the AllowSelfSignedServerCert

attribute to 1.4. To allow the common name of a CA-issued SSL certificate to not match the host name

of the Spark server, set the CAIssuedCertNamesMismatch attribute to 1.5. Choose one:

l To configure the driver to load SSL certificates from a specific PEM file when

Architecting the Future of Big Data

Hortonworks Inc. Page 40

verifying the server, set the TrustedCerts attribute to the full path of thePEM file.

l To use the trusted CA certificates PEM file that is installed with the driver, do notspecify a value for the TrustedCerts attribute.

6. To configure two-way SSL verification, set the TwoWaySSL attribute to 1 and then dothe following:

a) Set the ClientCert attribute to the full path of the PEM file containing the client'scertificate.

b) Set the ClientPrivateKey attribute to the full path of the file containing theclient's private key.

c) If the private key file is protected with a password, set theClientPrivateKeyPassword attribute to the password.

7. Save the odbc.ini configuration file.

Testing the ConnectionTo test the connection, you can use an ODBC-enabled client application. For a basicconnection test, you can also use the test utilities that are packaged with your drivermanager installation. For example, the iODBC driver manager includes simple utilities callediodbctest and iodbctestw. Similarly, the unixODBC driver manager includes simple utilitiescalled isql and iusql.

Using the iODBC Driver Manager

You can use the iodbctest and iodbctestw utilities to establish a test connection with yourdriver. Use iodbctest to test how your driver works with an ANSI application, or useiodbctestw to test how your driver works with a Unicode application.

Note: There are 32-bit and 64-bit installations of the iODBC driver manageravailable. If you have only one or the other installed, then the appropriate version ofiodbctest (or iodbctestw) is available. However, if you have both 32- and 64-bitversions installed, then you need to ensure that you are running the version from thecorrect installation directory.

For more information about using the iODBC driver manager, see http://www.iodbc.org.

To test your connection using the iODBC driver manager:1. Run iodbctest or iodbctestw.2. Optionally, if you do not remember the DSN, then type a question mark (?) to see a list

of available DSNs.3. Type an ODBC connection string using the following format, specifying additional

connection attributes as needed:DSN=DataSourceName;Key=Value

Architecting the Future of Big Data

Hortonworks Inc. Page 41

DataSourceName is the DSN that you are using for the connection. Key is anyconnection attribute that is not already specified as a configuration key in the DSN,and Value is the value for the attribute. Add key-value pairs to the connection string asneeded, separating each pair with a semicolon (;).Or, If you are using a DSN-less connection, then type an ODBC connection stringusing the following format, specifying additional connection attributes as needed:Driver=DriverNameOrFile;HOST=MySparkServer;PORT=PortNumber;Schema=DefaultSchema;SparkServerType=ServerType

The placeholders in the connection string are defined as follows:l DriverNameOrFile is either the symbolic name of the installed driver defined inthe odbcinst.ini file or the absolute path of the shared object file for the driver. Ifyou use the symbolic name, then you must ensure that the odbcinst.ini file is con-figured to point the symbolic name to the shared object file. For more inform-ation, see "Specifying ODBC drivers in odbcinst.ini" on page 34.

l MySparkServer is the IP address or hostname of the Spark Server.l PortNumber is the number of the TCP port that the Spark server uses to listenfor client connections.

l DefaultSchema is the database schema to use when a schema is not explicitlyspecified in a query.

l ServerType is 1 (for Shark Server), 2 (for Shark Server 2) or 3 (for Spark ThriftServer).

If the connection is successful, then the SQL> prompt appears.

Note: For information about the connection attributes that are available, see "DriverConfiguration Options" on page 55.

Using the unixODBC Driver Manager

You can use the isql and iusql utilities to establish a test connection with your driver and yourDSN. isql and iusql can only be used to test connections that use a DSN. Use isql to testhow your driver works with an ANSI application, or use iusql to test how your driver workswith a Unicode application.

Note: There are 32-bit and 64-bit installations of the unixODBC driver manageravailable. If you have only one or the other installed, then the appropriate version ofisql (or iusql) is available. However, if you have both 32- and 64-bit versionsinstalled, then you need to ensure that you are running the version from the correctinstallation directory.

For more information about using the unixODBC driver manager, seehttp://www.unixodbc.org.

Architecting the Future of Big Data

Hortonworks Inc. Page 42

To test your connection using the unixODBC driver manager:Run either isql or iusql using the following syntax:isql DataSourceName

iusql DataSourceName

DataSourceName is the DSN that you are using for the connection.

If the connection is successful, then the SQL> prompt appears.

Note: For information about the available options, run isql or iusql without providinga DSN.

Configuring Logging OptionsTo help troubleshoot issues, you can enable logging in the driver.

Important: Only enable logging long enough to capture an issue. Loggingdecreases performance and can consume a large quantity of disk space.

Use the LogLevel key to set the amount of detail included in log files. Table 2 lists thelogging levels provided by the Hortonworks ODBC Driver with SQL Connector for ApacheSpark, in order from least verbose to most verbose.

LogLevel Value Description

0 Disables all logging.

1 Logs very severe error events that will lead the driver to abort.

2 Logs error events that might still allow the driver to continuerunning.

3 Logs potentially harmful situations.

4 Logs general information that describes the progress of thedriver.

5 Logs detailed information that is useful for debugging the driver.

6 Logsmore detailed information than LogLevel=5

Table 2. Hortonworks ODBC Driver with SQL Connector for Apache Spark Logging Levels

To enable logging:1. Open the hortonworks.sparkodbc.ini configuration file in a text editor.

Architecting the Future of Big Data

Hortonworks Inc. Page 43

2. Set the LogLevel key to the desired level of information to include in log files. Forexample:LogLevel=2

3. Set the LogPath key to the full path to the folder where you want to save log files. Forexample:LogPath=/localhome/employee/Documents

4. Set the LogFileCount key to the maximum number of log files to keep.Note: After the maximum number of log files is reached, each time an additional fileis created, the driver deletes the oldest log file.

5. Set the LogFileSize key to the maximum size of each log file in megabytes (MB).Note: After the maximum file size is reached, the driver creates a new file andcontinues logging.

6. Save the hortonworks.sparkodbc.ini configuration file.7. Restart your ODBC application to ensure that the new settings take effect.

The Hortonworks ODBC Driver with SQL Connector for Apache Spark produces a log filenamed SparkODBC_driver.log at the location you specify using the LogPath key.

To disable logging:1. Open the hortonworks.sparkodbc.ini configuration file in a text editor.2. Set the LogLevel key to 0.3. Save the hortonworks.sparkodbc.ini configuration file.

Architecting the Future of Big Data

Hortonworks Inc. Page 44

Authentication OptionsTo connect to a Spark server, you must configure the Hortonworks ODBC Driver with SQLConnector for Apache Spark to use the authentication mechanism that matches the accessrequirements of the server and provides the necessary credentials. To determine theauthentication settings that your Spark server requires, check the server type and then referto the corresponding section below.

Shark ServerYoumust use No Authentication as the authentication mechanism. Shark Server instancesdo not support authentication.

Shark Server 2 or Spark Thrift Server on an HDInsight Dis-tributionIf you are connecting to HDInsight Emulator running on Windows Azure, then you must usetheWindows Azure HDInsight Emulator mechanism.

If you are connecting to HDInsight Service running on Windows Azure, then you must usetheWindows Azure HDInsight Service mechanism.

Shark Server 2 or Spark Thrift Server on a non-HDInsightDistributionNote: Most default configurations of Shark Server 2 or Spark Thrift Server on non-HDInsight Distributions require User Name authentication.

Configuring authentication for a connection to a Shark Server 2 or Spark Thrift Serverinstance on a non-HDInsight Distribution involves setting the authentication mechanism, theThrift transport protocol, and SSL support. To determine the settings that you need to use,check the following three properties in the hive-site.xml file in the Spark server that you areconnecting to:

l hive.server2.authenticationl hive.server2.transport.model hive.server2.use.SSL

Use Table 3 to determine the authentication mechanism that you need to configure, basedon the hive.server2.authentication value in the hive-site.xml file:

Architecting the Future of Big Data

Hortonworks Inc. Page 45

hive.server2.authentication Authentication Mechanism

NOSASL No Authentication

KERBEROS Kerberos

NONE User Name

LDAP User Name and Password

Table 3. Authentication Mechanism to Use

Use Table 4 to determine the Thrift transport protocol that you need to configure, based onthe hive.server2.authentication and hive.server2.transport.mode values in the hive-site.xmlfile:

hive.server2.authentication hive.server2.transport.mode Thrift TransportProtocol

NOSASL binary Binary

KERBEROS binary or http SASL or HTTP

NONE binary or http SASL or HTTP

LDAP binary or http SASL or HTTP

Table 4. Thrift Transport Protocol Setting to Use

To determine whether SSL should be enabled or disabled for your connection, check thehive.server2.use.SSL value in the hive-site.xml file. If the value is true, then you must enableand configure SSL in your connection. If the value is false, then you must disable SSL inyour connection.

For detailed instructions on how to configure authentication when using theWindows driver,see "Configuring Authentication" on page 10.

For detailed instructions on how to configure authentication when using a non-Windowsdriver, see "Configuring Authentication" on page 36.

Architecting the Future of Big Data

Hortonworks Inc. Page 46

Using a Connection StringFor some applications, you may need to use a connection string to connect to your datasource. For detailed information about how to use a connection string in an ODBCapplication, refer to the documentation for the application that you are using.

The connection strings in the following topics are examples showing the minimum set ofconnection attributes that you must specify to successfully connect to the data source.Depending on the configuration of the data source and the type of connection you areworking with, you may need to specify additional connection attributes. For detailedinformation about all the attributes that you can use in the connection string, see "DriverConfiguration Options" on page 55.

l "DSN Connection String Example" on page 46l "DSN-less Connection String Examples" on page 46

DSN Connection String ExampleThe following is an example of a connection string for a connection that uses a DSN:DSN=DataSourceName;

DataSourceName is the DSN that you are using for the connection.

You can set additional configuration options by appending key-value pairs to the connectionstring. Configuration options that are passed in using a connection string take precedenceover configuration options that are set in the DSN.

For information about creating a DSN on aWindows computer, see "Creating a DataSource Name" on page 7. For information about creating a DSN on a non-Windowscomputer, see "Defining DSNs in odbc.ini" on page 33.

DSN-less Connection String ExamplesSome applications provide support for connecting to a data source using a driver without aDSN. To connect to a data source without using a DSN, use a connection string instead.

The placeholders in the examples are defined as follows, in alphabetical order:l DomainName is the fully qualified domain name of the Spark server host.l PortNumber is the number of the TCP port that the Spark server uses to listen for cli-ent connections.

l Realm is the Kerberos realm of the Spark server host.l Server is the IP address or host name of the Spark server to which you are con-necting.

l ServerURL is the partial URL corresponding to the Spark server.

Architecting the Future of Big Data

Hortonworks Inc. Page 47

l ServiceName is the Kerberos service principal name of the Spark server.l YourPassword is the password corresponding to your user name.l YourUserName is the user name that you use to access the Spark server.

Connecting to a Shark Server Instance

The following is the format of a DSN-less connection string that connects to a Shark Serverinstance:Driver=Hortonworks Spark ODBC Driver;SparkServerType=1;Host=Server;Port=PortNumber;

For example:Driver=Hortonworks Spark ODBC Driver;SparkServerType=1;Host=192.168.222.160;Port=10000;

Connecting to a Standard Shark Server 2 Instance

The following is the format of a DSN-less connection string for a standard connection to aShark Server 2 instance. Most default configurations of Shark Server 2 require User Nameauthentication. When configured to provide User Name authentication, the driver usesanonymous as the user name by default.Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=Server;Port=PortNumber;AuthMech=2;

For example:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=192.168.222.160;Port=10000;AuthMech=2;

Connecting to a Standard Spark Thrift Server Instance

The following is the format of a DSN-less connection string for a standard connection to aSpark Thrift Server instance. By default, the driver is configured to connect to a Spark ThriftServer instance. Most default configurations of Spark Thrift Server require User Nameauthentication. When configured to provide User Name authentication, the driver usesanonymous as the user name by default.Driver=Hortonworks Spark ODBC Driver;Host=Server;Port=PortNumber;AuthMech=2;

For example:Driver=Hortonworks Spark ODBCDriver;Host=192.168.222.160;Port=10000;AuthMech=2;

Architecting the Future of Big Data

Hortonworks Inc. Page 48

Connecting to a Spark Thrift Server InstanceWithout Authentication

The following is the format of a DSN-less connection string that for a Spark Thrift Serverinstance that does not require authentication.Driver=Hortonworks Spark ODBCDriver;Host=Server;Port=PortNumber;AuthMech=0;

For example:Driver=Hortonworks Spark ODBCDriver;Host=192.168.222.160;Port=10000;AuthMech=0;

Connecting to a Spark Server that Requires Kerberos Authentication

The following is the format of a DSN-less connection string that connects to a Shark Server2 instance requiringKerberos authentication:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=Server;Port=PortNumber;AuthMech=1;KrbRealm=Realm;KrbHostFQDN=DomainName;KrbServiceName=ServiceName;

For example:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=192.168.222.160;Port=10000;AuthMech=1;KrbRealm=HORTONWORKS;KrbHostFQDN=localhost.localdomain;KrbServiceName=spark;

The following is the format of a DSN-less connection string that connects to a Spark ThriftServer instance requiringKerberos authentication. By default, the driver is configured toconnect to a Spark Thrift Server instance.Driver=Hortonworks Spark ODBC Driver;Host=Server;Port=PortNumber;AuthMech=1;KrbRealm=Realm;KrbHostFQDN=DomainName;KrbServiceName=ServiceName;

For example:Driver=Hortonworks Spark ODBCDriver;Host=192.168.222.160;Port=10000;AuthMech=1;KrbRealm=HORTONWORKS;KrbHostFQDN=localhost.localdomain;KrbServiceName=spark;

Connecting to a Spark Server that Requires User Name and Password Authentication

The following is the format of a DSN-less connection string that connects to a Shark Server2 instance requiringUser Name and Password authentication:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=Server;Port=PortNumber;AuthMech=3;UID=YourUserName;PWD=YourPassword;

For example:

Architecting the Future of Big Data

Hortonworks Inc. Page 49

Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=192.168.222.160;Port=10000;AuthMech=3;UID=hortonworks;PWD=hortonworks;

The following is the format of a DSN-less connection string that connects to a Spark ThriftServer instance requiringUser Name and Password authentication. By default, the driveris configured to connect to a Spark Thrift Server instance.Driver=Hortonworks Spark ODBC Driver;Host=Server;Port=PortNumber;AuthMech=3;UID=YourUserName;PWD=YourPassword;

For example:Driver=Hortonworks Spark ODBCDriver;Host=192.168.222.160;Port=10000;AuthMech=3;UID=hortonworks;PWD=hortonworks;

Connecting to a Spark Server onWindowsAzure HDInsight Emulator

The following is the format of a DSN-less connection string that connects to a Shark Server2 instance running on Windows Azure HDInsight Emulator:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=Server;Port=PortNumber;AuthMech=5;UID=YourUserName;PWD=YourPassword;HTTPPath=ServerURL;

For example:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=192.168.222.160;Port=10000;AuthMech=5;UID=hortonworks;PWD=hortonworks;HTTPPath=gateway/sandbox/spark;

The following is the format of a DSN-less connection string that connects to a Spark ThriftServer instance running on Windows Azure HDInsight Emulator. By default, the driver isconfigured to connect to a Spark Thrift Server instance.Driver=Hortonworks Spark ODBC Driver;Host=Server;Port=PortNumber;AuthMech=5;UID=YourUserName;PWD=YourPassword;HTTPPath=ServerURL;

For example:Driver=Hortonworks Spark ODBC Driver;Host=192.168.222.160;Port=10000;AuthMech=5;UID=hortonworks;PWD=hortonworks;HTTPPath=gateway/sandbox/spark;

Connecting to a Spark Server onWindowsAzure HDInsight Service

The following is the format of a DSN-less connection string that connects to a Shark Server2 instance running on Windows Azure HDInsight Service:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=Server;Port=PortNumber;AuthMech=6;UID=YourUserName;PWD=YourPassword;HTTPPath=ServerURL;

Architecting the Future of Big Data

Hortonworks Inc. Page 50

For example:Driver=Hortonworks Spark ODBC Driver;SparkServerType=2;Host=192.168.222.160;Port=10000;AuthMech=6;UID=hortonworks;PWD=hortonworks;HTTPPath=gateway/sandbox/spark;

The following is the format of a DSN-less connection string that connects to a Spark ThriftServer instance running on Windows Azure HDInsight Service. By default, the driver isconfigured to connect to a Spark Thrift Server instance.Driver=Hortonworks Spark ODBC Driver;Host=Server;Port=PortNumber;AuthMech=6;UID=YourUserName;PWD=YourPassword;HTTPPath=ServerURL;

For example:Driver=Hortonworks Spark ODBC Driver;Host=192.168.222.160;Port=10000;AuthMech=6;UID=hortonworks;PWD=hortonworks;HTTPPath=gateway/sandbox/spark;

Architecting the Future of Big Data

Hortonworks Inc. Page 51

FeaturesMore information is provided on the following features of the Hortonworks ODBC Driver withSQL Connector for Apache Spark:

l "SQLConnector for HiveQL" on page 51l "Data Types" on page 51l "Catalog and Schema Support" on page 52l "spark_system Table" on page 53l "Server-Side Properties" on page 53l "Get Tables With Query" on page 53l "Active Directory" on page 54l "Write-back" on page 54

SQL Connector for HiveQLThe native query language supported by Spark is HiveQL. For simple queries, HiveQL is asubset of SQL-92. However, the syntax is different enough that most applications do notworkwith native HiveQL.

To bridge the difference between SQL and HiveQL, the SQL Connector feature translatesstandard SQL-92 queries into equivalent HiveQL queries. The SQL Connector performssyntactical translations and structural transformations. For example:

l Quoted Identifiers: The double quotes (") that SQL uses to quote identifiers aretranslated into back quotes (`) to match HiveQL syntax. The SQL Connector needs tohandle this translation because evenwhen a driver reports the back quote as thequote character, some applications still generate double-quoted identifiers.

l Table Aliases: Support is provided for the AS keyword between a table referenceand its alias, which HiveQL normally does not support.

l JOIN, INNER JOIN , and CROSS JOIN : SQL JOIN, INNER JOIN, and CROSS JOINsyntax is translated to HiveQL JOIN syntax.

l TOP N/LIMIT: SQL TOP N queries are transformed to HiveQL LIMIT queries.

Data TypesThe Hortonworks ODBC Driver with SQL Connector for Apache Spark supports manycommon data formats, converting between Spark data types and SQL data types.

Table 5 lists the supported data type mappings.

Architecting the Future of Big Data

Hortonworks Inc. Page 52

Spark Type SQL Type

BIGINT SQL_BIGINT

BINARY SQL_VARBINARY

BOOLEAN SQL_BIT

CHAR(n) SQL_CHAR

DATE SQL_TYPE_DATE

DECIMAL SQL_DECIMAL

DECIMAL(p,s) SQL_DECIMAL

DOUBLE SQL_DOUBLE

FLOAT SQL_REAL

INT SQL_INTEGER

SMALLINT SQL_SMALLINT

STRING SQL_VARCHAR

TIMESTAMP SQL_TYPE_TIMESTAMP

TINYINT SQL_TINYINT

VARCHAR(n) SQL_VARCHAR

Table 5. Supported Data Types

Note:The aggregate types (ARRAY, MAP, and STRUCT) are not yet supported.Columns of aggregate types are treated as STRING columns.

Catalog and Schema SupportThe Hortonworks ODBC Driver with SQL Connector for Apache Spark supports bothcatalogs and schemas to make it easy for the driver to work with various ODBC applications.Since Spark only organizes tables into schemas/databases, the driver provides a syntheticcatalog called “SPARK” under which all of the schemas/databases are organized. Thedriver also maps the ODBC schema to the Spark schema/database.

Architecting the Future of Big Data

Hortonworks Inc. Page 53

The Hortonworks ODBC Driver with SQL Connector for Apache Spark supports queryingagainst Hive schemas.

spark_system TableA pseudo-table called spark_system can be used to query for Spark cluster systemenvironment information. The pseudo-table is under the pseudo-schema called spark_system. The table has two STRING type columns, envkey and envvalue. Standard SQL canbe executed against the spark_system table. For example:SELECT * FROM SPARK.spark_system.spark_system WHERE envkey LIKE'%spark%'

The above query returns all of the Spark system environment entries whose key containsthe word “spark.” A special query, set -v, is executed to fetch system environmentinformation. Some versions of Spark do not support this query. For versions of Spark that donot support querying system environment information, the driver returns an empty result set.

Server-Side PropertiesThe Hortonworks ODBC Driver with SQL Connector for Apache Spark allows you to setserver-side properties via a DSN. Server-side properties specified in a DSN affect only theconnection that is established using the DSN.

You can also specify server-side properties for connections that do not use a DSN. To dothis, use the Hortonworks Spark ODBC Driver Configuration tool that is installed with theWindows version of the driver, or set the appropriate configuration options in yourconnection string or the hortonworks.sparkodbc.ini file. Properties specified in the driverconfiguration tool or the hortonworks.sparkodbc.ini file apply to all connections that use theHortonworks ODBC Driver with SQL Connector for Apache Spark.

For more information about setting server-side properties when using the Windows driver,see "Configuring Server-Side Properties" on page 17. For information about setting server-side properties when using the driver on a non-Windows platform, see "Driver ConfigurationOptions" on page 55.

Get Tables With QueryThe Get TablesWith Query configuration option allows you to choose whether to use theSHOWTABLES query or the GetTables API call to retrieve table names from a database.

Shark Server 2 and Spark Thrift Server both have a limit on the number of tables that can bein a database when handling the GetTables API call. When the number of tables in adatabase is above the limit, the API call will return a stack overflow error or a timeout error.The exact limit and the error that appears depends on the JVM settings.

As a workaround for this issue, enable the Get Tables with Query configuration option (orGetTablesWithQuery key) to use the query instead of the API call.

Architecting the Future of Big Data

Hortonworks Inc. Page 54

Active DirectoryThe Hortonworks ODBC Driver with SQL Connector for Apache Spark supports ActiveDirectory Kerberos on Windows. There are two prerequisites for using Active DirectoryKerberos on Windows:

l MIT Kerberos is not installed on the client Windows machine.l The MIT Kerberos Hadoop realm has been configured to trust the Active Directoryrealm so that users in the Active Directory realm can access services in the MITKerberos Hadoop realm. For more information, see Setting up One-Way Trust withActive Directory in the Hortonworks documentation athttp://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.1.7/bk_installing_manually_book/content/ch23s05.html.

Write-backThe Hortonworks ODBC Driver with SQL Connector for Apache Spark supports translationfor INSERT, UPDATE, and DELETE syntax when connecting to a Shark Server 2 or SparkThrift Server instance.

Architecting the Future of Big Data

Hortonworks Inc. Page 55

Driver Configuration OptionsDriver Configuration Options lists the configuration options available in the HortonworksODBC Driver with SQL Connector for Apache Spark alphabetically by field or button label.Options having only key names, that is, not appearing in the user interface of the driver, arelisted alphabetically by key name.

When creating or configuring a connection from a Windows computer, the fields and buttonsare available in the Hortonworks Spark ODBC Driver Configuration tool and the followingdialog boxes:

l Hortonworks Spark ODBC Driver DSN Setupl HTTP Propertiesl SSL Optionsl Advanced Optionsl Server Side Properties

When using a connection string or configuring a connection from a Linux/Mac OS Xcomputer, use the key names provided.

Note: You can pass in configuration options in your connection string, or set them inyour odbc.ini and hortonworks.sparkodbc.ini files if you are using a non-Windowsversion of the driver. Configuration options set in a hortonworks.sparkodbc.ini fileapply to all connections, whereas configuration options passed in in the connectionstring or set in an odbc.ini file are specific to a connection. Configuration optionspassed in using the connection string take precedence over configuration optionsset in odbc.ini. Configuration options set in odbc.ini take precedence overconfiguration options set in hortonworks.sparkodbc.ini

Configuration Options Appearing in the User InterfaceThe following configuration options are accessible via the Windows user interface for theHortonworks ODBC Driver with SQL Connector for Apache Spark, or via the key namewhen using a connection string or configuring a connection from a Linux/Mac OS Xcomputer:

l "Allow Common Name Host NameMismatch" on page 56

l "Allow Self-signed Server Cer-tificate" on page 57

l "Apply Properties with Queries" onpage 57

l "Async Exec Poll Interval" on page57

l "Host FQDN" on page 62l "HTTP Path" on page 63l Invalid Session Auto Recover onpage 1

l "Mechanism" on page 63l "Password" on page 64l "Port" on page 64

Architecting the Future of Big Data

Hortonworks Inc. Page 56

l "Binary Column Length" on page 58l "Client Certificate File" on page 58l "Client Private Key File" on page 58l "Client Private Key Password" onpage 59

l "Convert Key Name to Lower Case"on page 59

l "Database" on page 59l "Decimal Column Scale" on page 60l "Default String Column Length" onpage 60

l "Delegation UID" on page 60l "Driver Config Take Precedence" onpage 60

l "Enable SSL" on page 61l "Fast SQLPrepare" on page 61l "Get Tables With Query" on page 61l "Host" on page 62

l "Realm" on page 64l "Rows Fetched Per Block" on page65

l "Save Password (Encrypted)" onpage 65

l "Service Name" on page 65l "Show System Table" on page 66l "Socket Timeout" on page 66l "Spark Server Type" on page 66l "Thrift Transport" on page 67l "Trusted Certificates" on page 67l "TwoWay SSL" on page 68l "Unicode SQL Character Types" onpage 68

l "Use Async Exec" on page 69l "UseNative Query" on page 69l "UseOnly SSPI Plugin" on page 69l "User Name" on page 70

Allow Common Name Host Name Mismatch

Key Name Default Value Required

CAIssuedCertNamesMismatch Clear (0) No

Description

When this option is enabled (1), the driver allows a CA-issued SSL certificate name to notmatch the host name of the Spark server.

When this option is disabled (0), the CA-issued SSL certificate namemust match the hostname of the Spark server.

Note: This setting is applicable only when SSL is enabled.

Note: This option is applicable only to the User Name and Password (SSL) andHTTPS authentication mechanisms.

Architecting the Future of Big Data

Hortonworks Inc. Page 57

Allow Self-signed Server Certificate

Key Name Default Value Required

AllowSelfSignedServerCert Clear (0) No

Description

When this option is enabled (1), the driver authenticates the Spark server even if the serveris using a self-signed certificate.

When this option is disabled (0), the driver does not allow self-signed certificates from theserver.

Note: This setting is applicable only when SSL is enabled.

Apply Properties with Queries

Key Name Default Value Required

ApplySSPWithQueries Selected (1) No

Description

When this option is enabled (1), the driver applies each server-side property by executing aset SSPKey=SSPValue query when opening a session to the Spark server.

When this option is disabled (0), the driver uses a more efficient method for applying server-side properties that does not involve additional network round-tripping. However, someShark Server 2 or Spark Thrift Server builds are not compatible with the more efficientmethod.

Note: When connecting to a Shark Server instance, ApplySSPWithQueries isalways enabled.

Async Exec Poll Interval

Key Name Default Value Required

AsyncExecPollInterval 100 No

Description

The time in milliseconds between each poll for the query execution status.

Architecting the Future of Big Data

Hortonworks Inc. Page 58

“Asynchronous execution” refers to the fact that the RPC call used to execute a queryagainst Spark is asynchronous. It does not mean that ODBC asynchronous operations aresupported.

Note: This option is applicable only to HDInsight clusters.

Binary Column Length

Key Name Default Value Required

BinaryColumnLength 32767 No

Description

The maximum data length for BINARY columns.

By default, the columns metadata for Spark does not specify a maximum data length forBINARY columns.

Client Certificate File

Key Name Default Value Required

ClientCert None Yes, if two-waySSL verification is enabled.

Description

The full path to the PEM file containing the client's SSL certificate.

Note: This setting is applicable only when two-way SSL is enabled.

Client Private Key File

Key Name Default Value Required

ClientPrivateKey None Yes, if two-waySSL verification is enabled.

Description

The full path to the PEM file containing the client's SSL private key.

If the private key file is protected with a password, then provide the password using thedriver configuration option "Client Private Key Password" on page 59.

Architecting the Future of Big Data

Hortonworks Inc. Page 59

Note: This setting is applicable only when two-way SSL is enabled.

Client Private Key Password

Key Name Default Value Required

ClientPrivateKeyPassword None Yes, if two-waySSL verification is enabledand the client's private keyfile is protected with apassword.

Description

The password of the private key file that is specified in the Client Private Key File field (theClientPrivateKey key).

Convert Key Name to Lower Case

Key Name Default Value Required

LCaseSspKeyName Selected (1) No

Description

When this option is enabled (1), the driver converts server-side property key names to alllower case characters.

When this option is disabled (0), the driver does not modify the server-side property keynames.

Database

Key Name Default Value Required

Schema default No

Description

The name of the database schema to use when a schema is not explicitly specified in aquery. You can still issue queries on other schemas by explicitly specifying the schema in thequery.

Note: To inspect your databases and determine the appropriate schema to use,type the show databases command at the Spark command prompt.

Architecting the Future of Big Data

Hortonworks Inc. Page 60

Decimal Column Scale

Key Name Default Value Required

DecimalColumnScale 10 No

Description

The maximum number of digits to the right of the decimal point for numeric data types.

Default String Column Length

Key Name Default Value Required

DefaultStringColumnLength 255 No

Description

The maximum number of characters that can be contained in STRING columns.

By default, the columns metadata for Spark does not specify a maximum length for STRINGcolumns.

Delegation UID

Key Name Default Value Required

DelegationUID None No

Description

Use this option to delegate all operations against Spark to a user that is different than theauthenticated user for the connection.

Note: This option is applicable onlywhen connecting to a Shark Server 2 and SparkThrift Server instance that supports this feature.

Driver Config Take Precedence

Key Name Default Value Required

DriverConfigTakePrecedence Clear (0) No

Architecting the Future of Big Data

Hortonworks Inc. Page 61

Description

When this option is enabled (1), driver-wide configurations take precedence over connectionand DSN settings.

When this option is disabled (0), connection and DSN settings take precedence instead.

Enable SSL

Key Name Default Value Required

Clear (0) No

Description

When this option is enabled (1), the client verifies the Spark server using SSL.

When this option is disabled (0), SSL is disabled.

Note:This option is applicable only when connecting to a Spark server that supports SSL.

Fast SQLPrepare

Key Name Default Value Required

FastSQLPrepare Clear (0) No

Description

When this option is enabled (1), the driver defers query execution to SQLExecute.

When this option is disabled (0), the driver does not defer query execution to SQLExecute.

When using Native Query mode, the driver will execute the HiveQL query to retrieve theresult set metadata for SQLPrepare. As a result, SQLPrepare might be slow. If the result setmetadata is not required after calling SQLPrepare, then enable Fast SQLPrepare.

Get Tables With Query

Key Name Default Value Required

GetTablesWithQuery Clear (0) No

Architecting the Future of Big Data

Hortonworks Inc. Page 62

Description

When this option is enabled (1), the driver uses the SHOW TABLES query to retrieve thenames of the tables in a database.

When this option is disabled (0), the driver uses the GetTables Thrift API call to retrieve thenames of the tables in a database.

Note: This option is applicable onlywhen connecting to a Shark Server 2 or SparkThrift Server instance.

Host

Key Name Default Value Required

HOST None Yes

Description

The IP address or host name of the Spark server.

Host FQDN

Key Name Default Value Required

KrbHostFQDN None Yes, if the authenticationmechanism is Kerberos.

Description

The fully qualified domain name of the Shark Server 2 or Spark Thrift Server host.

You can set the value of Host FQDN to _HOST to use the Spark server host name as thefully qualified domain name for Kerberos authentication.

Architecting the Future of Big Data

Hortonworks Inc. Page 63

HTTP Path

Key Name Default Value Required

HTTPPath /sparkhive2 if usingWindows Azure HDInsightService (6)

/ if using non-WindowsAzure HDInsight Servicewith Thrift Transport set toHTTP(2)

No

Description

The partial URL corresponding to the Spark server.

Mechanism

Key Name Default Value Required

AuthMech No Authentication (0) ifyou are connecting toSpark Server 1.

User Name (2) if you areconnecting to SparkServer 2.

No

Description

The authentication mechanism to use.

Select one of the following settings, or set the key to the corresponding number:l No Authentication (0)l Kerberos (1)l User Name (2)l User Name and Password (3)l Windows Azure HDInsight Emulator (5)l Windows Azure HDInsight Service (6)

Architecting the Future of Big Data

Hortonworks Inc. Page 64

Password

Key Name Default Value Required

PWD None Yes, if the authenticationmechanism is User Nameand Password,Windows

Azure HDInsightEmulator, orWindows

Azure HDInsightService.

Description

The password corresponding to the user name that you provided in the User Name field (theUID key).

Port

Key Name Default Value Required

PORT l non-HDInsightclusters: 10000

l Windows AzureHDInsight Emu-lator: 10001

l Windows AzureHDInsight Service:443

Yes

Description

The number of the TCP port that the Spark server uses to listen for client connections.

Realm

Key Name Default Value Required

KrbRealm Depends on your Kerberosconfiguration.

No

Architecting the Future of Big Data

Hortonworks Inc. Page 65

Description

The realm of the Shark Server 2 or Spark Thrift Server host.

If your Kerberos configuration already defines the realm of the Shark Server 2 or SparkThrift Server host as the default realm, then you do not need to configure this option.

Rows Fetched Per Block

Key Name Default Value Required

RowsFetchedPerBlock 10000 No

Description

The maximum number of rows that a query returns at a time.

Any positive 32-bit integer is a valid value, but testing has shown that performance gains aremarginal beyond the default value of 10000 rows.

Save Password (Encrypted)

Key Name Default Value Required

N/A Clear (0) No

Description

This option is available only in theWindows driver. It appears in the Hortonworks SparkODBC Driver DSN Setup dialog box and the SSL Options dialog box.

When this option is enabled (1), the specified password is saved in the registry.

When this option is disabled (0), the specified password is not saved in the registry.

Important: The password is obscured (not saved in plain text). However, it is stillpossible for the encrypted password to be copied and used.

Service Name

Key Name Default Value Required

Yes, if the authenticationmechanism is Kerberos.

Architecting the Future of Big Data

Hortonworks Inc. Page 66

Description

The Kerberos service principal name of the Spark server.

Show System Table

Key Name Default Value Required

ShowSystemTable Clear (0) No

Description

When this option is enabled (1), the driver returns the spark_system table for catalogfunction calls such as SQLTables and SQLColumns

When this option is disabled (0), the driver does not return the spark_system table forcatalog function calls.

Socket Timeout

Key Name Default Value Required

SocketTimeout 30 No

Description

The number of seconds that an operation can remain idle before it is closed.

Note: This option is applicable onlywhen asynchronous query execution is beingused against Shark Server 2 or Spark Thrift Server instances.

Spark Server Type

Key Name Default Value Required

SparkServerType Spark Thrift Server (3) No

Description

Select Shark Server or set the key to 1 if you are connecting to a Shark Server instance.

Select Shark Server 2 or set the key to 2 if you are connecting to a Shark Server 2instance.

Architecting the Future of Big Data

Hortonworks Inc. Page 67

Select Spark Thrift Server or set the key to 3 if you are connecting to a Spark Thrift Serverinstance.

Thrift Transport

Key Name Default Value Required

ThriftTransport Binary (0) if you areconnecting to SparkServer 1.

SASL (1) if you areconnecting to SparkServer 2.

No

Description

The transport protocol to use in the Thrift layer.

Select one of the following settings, or set the key to the corresponding number:l Binary (0)l SASL (1)l HTTP (2)

Trusted Certificates

Key Name Default Value Required

TrustedCerts The cacerts.pem file in thelib folder or subfolderwithin the driver'sinstallation directory.

The exact file path variesdepending on the versionof the driver that isinstalled. For example, thepath for theWindowsdriver is different from thepath for the MacOS X driver.

No

Architecting the Future of Big Data

Hortonworks Inc. Page 68

Description

The location of the PEM file containing trusted CA certificates for authenticating the Sparkserver when using SSL.

If this option is not set, then the driver will default to using the trusted CA certificatesPEMfile installed by the driver.

Note: This setting is applicable only when SSL is enabled.

Two Way SSL

Key Name Default Value Required

TwoWaySSL Clear (0) No

Description

When this option is enabled (1), the client and the Spark server verify each other using SSL.See also the driver configuration options "Client Certificate File" on page 58, "Client PrivateKey File" on page 58, and "Client Private Key Password" on page 59.

When this option is disabled (0), then the server does not verify the client. Depending onwhether one-way SSL is enabled, the client may verify the server. For more information, see"Enable SSL" on page 61.

Note: This option is applicable onlywhen connecting to a Spark server thatsupports SSL. You must enable SSL before TwoWay SSL can be configured. Formore information, see "Enable SSL" on page 61.

Unicode SQL Character Types

Key Name Default Value Required

UseUnicodeSqlCharacterTypes Clear (0) No

Description

When this option is enabled (1), the driver returns SQL_WVARCHAR for STRING andVARCHAR columns, and returns SQL_WCHAR for CHAR columns.

When this option is disabled (0), the driver returns SQL_VARCHAR for STRING andVARCHAR columns, and returns SQL_CHAR for CHAR columns.

Architecting the Future of Big Data

Hortonworks Inc. Page 69

Use Async Exec

Key Name Default Value Required

EnableAsyncExec Clear (0) No

Description

When this option is enabled (1), the driver uses an asynchronous version of the API callagainst Spark for executing a query.

When this option is disabled (0), the driver executes queries synchronously.

Use Native Query

Key Name Default Value Required

UseNativeQuery Clear (0) No

Description

When this option is enabled (1), the driver does not transform the queries emitted by anapplication, and executes HiveQL queries directly.

When this option is disabled (0), the driver transforms the queries emitted by an applicationand converts them into an equivalent form in HiveQL.

Note: If the application is Spark-aware and already emits HiveQL, then enable thisoption to avoid the extra overhead of query transformation.

Use Only SSPI Plugin

Key Name Default Value Required

UseOnlySSPI Clear (0) No

Description

When this option is enabled (1), the driver handles Kerberos authentication by using theSSPI plugin instead of MIT Kerberos by default.

When this option is disabled (0), the driver uses MIT Kerberos to handle Kerberosauthentication, and only uses the SSPI plugin if the gssapi library is not available.

Important: This option is available only in the Windows driver.

Architecting the Future of Big Data

Hortonworks Inc. Page 70

User Name

Key Name Default Value Required

UID For User Nameauthentication only, thedefault value isanonymous

No, if the authenticationmechanism is User Name.

Yes, if the authenticationmechanism is User Nameand Password,WindowsAzure HDInsightEmulator, orWindowsAzure HDInsightService.

Description

The user name that you use to access Shark Server 2 or Spark Thrift Server.

Configuration Options Having Only Key NamesThe following configuration options do not appear in the Windows user interface for theHortonworks ODBC Driver with SQL Connector for Apache Spark and are only accessiblewhen using a connection string or configuring a connection from a Linux/Mac OS Xcomputer:

l "ADUserNameCase" on page 70l "Driver" on page 71l "ForceSynchronousExec" on page 71l "http.header." on page 72l "SSP_" on page 72

ADUserNameCase

Default Value Required

Unchanged No

Architecting the Future of Big Data

Hortonworks Inc. Page 71

Description

Use this option to control whether the driver changes the user name part of an AD KerberosUPN to all upper case or all lower case. The following values are possible:

l Upper: Change the user name to all upper case.l Lower: Change the user name to all lower case.l Unchanged: Do not modify the user name.

Note: This option is applicable onlywhen using Active Directory Kerberos from aWindows client machine to authenticate.

Driver

Default Value Required

The default value varies depending on theversion of the driver that is installed. Forexample, the value for the Windows driveris different from the value of the MacOS X driver.

Yes

Description

The name of the installed driver (Hortonworks Spark ODBC Driver) or the absolute path ofthe Hortonworks ODBC Driver with SQL Connector for Apache Spark shared object file.

ForceSynchronousExec

Default Value Required

0 No

Description

When this option is enabled (1), the driver is forced to execute queries synchronously whenconnected to an HDInsight cluster.

When this option is disabled (0), the driver is able to execute queries asynchronously whenconnected to an HDInsight cluster.

Note: This option is applicable only to HDInsight clusters.

Architecting the Future of Big Data

Hortonworks Inc. Page 72

http.header.

Default Value Required

None No

Description

Set a customHTTP header by using the following syntax, where HeaderKey is the name ofthe header to set and HeaderValue is the value to assign to the header:http.header.HeaderKey=HeaderValue

For example:http.header.AUTHENTICATED_USER=john

After the driver applies the header, the http.header. prefix is removed from the DSN entry,leaving an entry of HeaderKey=HeaderValue

The example above would create the following custom HTTP header:AUTHENTICATED_USER: john

Note:The http.header. prefix is case-sensitive. This option is applicable only when youare using HTTP as the Thrift transport protocol. For more information, see "ThriftTransport" on page 67.

SSP_

Default Value Required

None No

Description

Set a server-side property by using the following syntax, where SSPKey is the name of theserver-side property to set and SSPValue is the value to assign to the server-side property:SSP_SSPKey=SSPValue

For example:SSP_mapred.queue.names=myQueue

After the driver applies the server-side property, the SSP_ prefix is removed from the DSNentry, leaving an entry of SSPKey=SSPValue

Note: The SSP_ prefix must be upper case.

Architecting the Future of Big Data

Hortonworks Inc. Page 73

Contact UsIf you have difficulty using the Hortonworks ODBC Driver with SQL Connector for ApacheSpark, please contact our support staff. We welcome your questions, comments, andfeature requests.

Please have a detailed summary of the client and server environment (OS version, patchlevel, Hadoop distribution version, Spark version, configuration, etc.) ready, before you callor write us. Supplying this information accelerates support.

By telephone:

USA: (855) 8-HORTON

International: (408) 916-4121

On the Internet:

Visit us at www.hortonworks.com

Architecting the Future of Big Data

Hortonworks Inc. Page 74

Third Party Trademarks and LicensesThird Party TrademarksKerberos is a trademark of the Massachusetts Institute of Technology (MIT).

Linux is the registered trademark of Linus Torvalds in Canada, United States and/or othercountries.

Microsoft, MSDN, Windows, Windows Azure, Windows Server, Windows Vista, and theWindows start button are trademarks or registered trademarks of Microsoft Corporation orits subsidiaries in Canada, United States and/or other countries.

Red Hat, Red Hat Enterprise Linux, and CentOS are trademarks or registered trademarksof Red Hat, Inc. or its subsidiaries in Canada, United States and/or other countries.

Apache Spark, Apache, and Spark are trademarks or registered trademarks of The ApacheSoftware Foundation or its subsidiaries in Canada, United States and/or other countries.

All other trademarks are trademarks of their respective owners.

Third Party LicensesBoost Software License - Version 1.0 - August 17th, 2003

Permission is hereby granted, free of charge, to any person or organization obtaining a copyof the software and accompanying documentation covered by this license (the "Software")to use, reproduce, display, distribute, execute, and transmit the Software, and to preparederivative works of the Software, and to permit third-parties to whom the Software isfurnished to do so, all subject to the following:

The copyright notices in the Software and this entire statement, including the above licensegrant, this restriction and the following disclaimer, must be included in all copies of theSoftware, in whole or in part, and all derivative works of the Software, unless such copies orderivative works are solely in the form of machine-executable object code generated by asource language processor.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OFMERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE AND NON-INFRINGEMENT. IN NO EVENT SHALL THE COPYRIGHT HOLDERS OR ANYONEDISTRIBUTINGTHE SOFTWARE BE LIABLE FOR ANY DAMAGES OR OTHERLIABILITY, WHETHER IN CONTRACT, TORTOR OTHERWISE, ARISING FROM, OUTOF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGSIN THE SOFTWARE.

Architecting the Future of Big Data

Hortonworks Inc. Page 75

cURL

COPYRIGHT AND PERMISSION NOTICE

Copyright (c) 1996 - 2015, Daniel Stenberg, [email protected].

All rights reserved.

Permission to use, copy, modify, and distribute this software for any purpose with or withoutfee is hereby granted, provided that the above copyright notice and this permission noticeappear in all copies.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OFMERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE ANDNONINFRINGEMENT OF THIRD PARTY RIGHTS. IN NO EVENT SHALL THEAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OROTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OROTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWAREOR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

Except as contained in this notice, the name of a copyright holder shall not be used inadvertising or otherwise to promote the sale, use or other dealings in this Software withoutprior written authorization of the copyright holder.

Cyrus SASL

Copyright (c) 1994-2012 Carnegie Mellon University. All rights reserved.

Redistribution and use in source and binary forms, with or without modification, arepermitted provided that the following conditions are met:1. Redistributions of source code must retain the above copyright notice, this list of

conditions and the following disclaimer.2. Redistributions in binary form must reproduce the above copyright notice, this list of

conditions and the following disclaimer in the documentation and/or other materialsprovided with the distribution.

3. The name "Carnegie Mellon University" must not be used to endorse or promoteproducts derived from this software without prior written permission. For permissionor any other legal details, please contactOffice of Technology TransferCarnegie Mellon University5000 Forbes AvenuePittsburgh, PA 15213-3890(412) 268-4387, fax: (412) [email protected]

4. Redistributions of any form whatsoever must retain the following acknowledgment:

Architecting the Future of Big Data

Hortonworks Inc. Page 76

"This product includes software developed by Computing Services at Carnegie MellonUniversity (http://www.cmu.edu/computing/)."

CARNEGIE MELLON UNIVERSITY DISCLAIMS ALL WARRANTIES WITH REGARD TOTHIS SOFTWARE, INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITYAND FITNESS, IN NO EVENT SHALL CARNEGIE MELLON UNIVERSITY BE LIABLEFOR ANY SPECIAL, INDIRECT OR CONSEQUENTIAL DAMAGES OR ANY DAMAGESWHATSOEVER RESULTINGFROM LOSS OF USE, DATA OR PROFITS, WHETHER INAN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISINGOUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THISSOFTWARE.

dtoa License

The author of this software is David M. Gay.

Copyright (c) 1991, 2000, 2001 by Lucent Technologies.

Permission to use, copy, modify, and distribute this software for any purpose without fee ishereby granted, provided that this entire notice is included in all copies of any softwarewhich is or includes a copy or modification of this software and in all copies of the supportingdocumentation for such software.

THIS SOFTWARE IS BEING PROVIDED "AS IS", WITHOUT ANY EXPRESS ORIMPLIED WARRANTY. IN PARTICULAR, NEITHER THE AUTHOR NOR LUCENTMAKES ANY REPRESENTATION OR WARRANTY OF ANY KIND CONCERNING THEMERCHANTABILITY OF THIS SOFTWARE OR ITS FITNESS FOR ANY PARTICULARPURPOSE.

Expat

Copyright (c) 1998, 1999, 2000 Thai Open Source Software Center Ltd

Permission is hereby granted, free of charge, to any person obtaining a copy of this softwareand associated documentation files (the "Software"), to deal in the Software withoutrestriction, including without limitation the rights to use, copy, modify, merge, publish,distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom theSoftware is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies orsubstantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OFMERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE ANDNOINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHTHOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHERIN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR

Architecting the Future of Big Data

Hortonworks Inc. Page 77

IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THESOFTWARE.

ICU License - ICU 1.8.1 and later

COPYRIGHT AND PERMISSION NOTICE

Copyright (c) 1995-2014 International Business Machines Corporation and others

All rights reserved.

Permission is hereby granted, free of charge, to any person obtaining a copy of this softwareand associated documentation files (the "Software"), to deal in the Software withoutrestriction, including without limitation the rights to use, copy, modify, merge, publish,distribute, and/or sell copies of the Software, and to permit persons to whom the Software isfurnished to do so, provided that the above copyright notice(s) and this permission noticeappear in all copies of the Software and that both the above copyright notice(s) and thispermission notice appear in supporting documentation.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OFMERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE ANDNONINFRINGEMENT OF THIRD PARTY RIGHTS. IN NO EVENT SHALL THECOPYRIGHT HOLDER OR HOLDERS INCLUDED IN THIS NOTICE BE LIABLE FORANY CLAIM, OR ANY SPECIAL INDIRECT OR CONSEQUENTIAL DAMAGES, OR ANYDAMAGES WHATSOEVER RESULTINGFROM LOSS OF USE, DATA OR PROFITS,WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUSACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCEOF THIS SOFTWARE.

Except as contained in this notice, the name of a copyright holder shall not be used inadvertising or otherwise to promote the sale, use or other dealings in this Software withoutprior written authorization of the copyright holder.

All trademarks and registered trademarks mentioned herein are the property of theirrespective owners.

OpenSSL

Copyright (c) 1998-2011 The OpenSSL Project. All rights reserved.

Redistribution and use in source and binary forms, with or without modification, arepermitted provided that the following conditions are met:1. Redistributions of source code must retain the above copyright notice, this list of

conditions and the following disclaimer.2. Redistributions in binary form must reproduce the above copyright notice, this list of

conditions and the following disclaimer in the documentation and/or other materialsprovided with the distribution.

Architecting the Future of Big Data

Hortonworks Inc. Page 78

3. All advertising materialsmentioning features or use of this software must display thefollowing acknowledgment:

"This product includes software developed by the OpenSSL Project for use in theOpenSSL Toolkit. (http://www.openssl.org/)"

4. The names "OpenSSL Toolkit" and "OpenSSL Project" must not be used to endorseor promote products derived from this software without prior written permission. Forwritten permission, please contact [email protected].

5. Products derived from this software may not be called "OpenSSL" nor may"OpenSSL" appear in their names without prior written permission of the OpenSSLProject.

6. Redistributions of any form whatsoever must retain the following acknowledgment:

"This product includes software developed by the OpenSSL Project for use in theOpenSSL Toolkit (http://www.openssl.org/)"

THIS SOFTWARE IS PROVIDED BY THE OpenSSL PROJECT "AS IS" AND ANYEXPRESSED OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THEIMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULARPURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE OpenSSL PROJECT OR ITSCONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL,EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO,PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, ORPROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANYTHEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THEUSE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCHDAMAGE.

This product includes cryptographic software written by Eric Young([email protected]).This product includes software written by Tim Hudson ([email protected]).

Original SSLeay License

Copyright (C) 1995-1998 Eric Young ([email protected])

All rights reserved.

This package is an SSL implementation written by Eric Young ([email protected]). Theimplementation was written so as to conform with Netscapes SSL.

This library is free for commercial and non-commercial use as long as the followingconditions are aheared to. The following conditions apply to all code found in thisdistribution, be it the RC4, RSA, lhash, DES, etc., code; not just the SSL code. The SSLdocumentation included with this distribution is covered by the same copyright terms exceptthat the holder is Tim Hudson ([email protected]).

Architecting the Future of Big Data

Hortonworks Inc. Page 79

Copyright remains Eric Young's, and as such any Copyright notices in the code are not to beremoved. If this package is used in a product, Eric Young should be given attribution as theauthor of the parts of the library used. This can be in the form of a textual message atprogram startup or in documentation (online or textual) provided with the package.

Redistribution and use in source and binary forms, with or without modification, arepermitted provided that the following conditions are met:1. Redistributions of source code must retain the copyright notice, this list of conditions

and the following disclaimer.2. Redistributions in binary form must reproduce the above copyright notice, this list of

conditions and the following disclaimer in the documentation and/or other materialsprovided with the distribution.

3. All advertising materialsmentioning features or use of this software must display thefollowing acknowledgement:

"This product includes cryptographic software written by Eric Young([email protected])"

The word 'cryptographic' can be left out if the rouines from the library being used arenot cryptographic related :-).

4. If you include any Windows specific code (or a derivative thereof) from the appsdirectory (application code) you must include an acknowledgement:

"This product includes software written by Tim Hudson ([email protected])"

THIS SOFTWARE IS PROVIDED BY ERIC YOUNG ``AS IS'' AND ANY EXPRESS ORIMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIEDWARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULARPURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE AUTHOR ORCONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL,EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO,PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, ORPROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANYTHEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THEUSE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCHDAMAGE.

The licence and distribution terms for any publically available version or derivative of thiscode cannot be changed. i.e. this code cannot simply be copied and put under anotherdistribution licence [including the GNU Public Licence.]

rapidjson

Tencent is pleased to support the open source community by making RapidJSON available.

Copyright (C) 2015 THL A29 Limited, a Tencent company, and Milo Yip. All rights reserved.

Architecting the Future of Big Data

Hortonworks Inc. Page 80

If you have downloaded a copy of the RapidJSON binary from Tencent, please note that theRapidJSON binary is licensed under the MIT License.

If you have downloaded a copy of the RapidJSON source code from Tencent, please notethat RapidJSON source code is licensed under the MIT License, except for the third-partycomponents listed below which are subject to different license terms. Your integration ofRapidJSON into your own projects may require compliance with the MIT License, as well asthe other licenses applicable to the third-party components included within RapidJSON.

A copy of the MIT License is included in this file.

Other dependencies and licenses:

Open Source Software Licensed Under the BSD License:

The msinttypes r29Copyright (c) 2006-2013 Alexander ChemerisAll rights reserved.

Redistribution and use in source and binary forms, with or without modification,are permitted provided that the following conditions are met:

l Redistributions of source code must retain the above copyright notice, thislist of conditions and the following disclaimer.

l Redistributions in binary form must reproduce the above copyright notice,this list of conditions and the following disclaimer in the documentationand/or other materials provided with the distribution.

l Neither the name of copyright holder nor the names of its contributorsmay be used to endorse or promote products derived from this softwarewithout specific prior written permission.

THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS``AS IS'' AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING,BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OFMERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE AREDISCLAIMED. IN NO EVENT SHALL THE REGENTS AND CONTRIBUTORSBE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL,EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOTLIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;LOSSOF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION)HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER INCONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OROTHERWISE) ARISING IN ANYWAY OUT OF THE USE OF THISSOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

Open Source Software Licensed Under the JSON License:

json.org

Architecting the Future of Big Data

Hortonworks Inc. Page 81

Copyright (c) 2002 JSON.orgAll Rights Reserved.

JSON_checkerCopyright (c) 2002 JSON.orgAll Rights Reserved.

Terms of the JSON License:

Permission is hereby granted, free of charge, to any person obtaining a copy ofthis software and associated documentation files (the "Software"), to deal in theSoftware without restriction, including without limitation the rights to use, copy,modify, merge, publish, distribute, sublicense, and/or sell copies of theSoftware, and to permit persons to whom the Software is furnished to do so,subject to the following conditions:

The above copyright notice and this permission notice shall be included in allcopies or substantial portions of the Software.

The Software shall be used for Good, not Evil.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANYKIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULARPURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THEAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM,DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OFCONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR INCONNECTION WITH THE SOFTWARE OR THE USE OR OTHERDEALINGS IN THE SOFTWARE.

Terms of the MIT License:

Permission is hereby granted, free of charge, to any person obtaining a copy ofthis software and associated documentation files (the "Software"), to deal in theSoftware without restriction, including without limitation the rights to use, copy,modify, merge, publish, distribute, sublicense, and/or sell copies of theSoftware, and to permit persons to whom the Software is furnished to do so,subject to the following conditions:

The above copyright notice and this permission notice shall be included in allcopies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANYKIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THEWARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULARPURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THEAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM,

Architecting the Future of Big Data

Hortonworks Inc. Page 82

DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OFCONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR INCONNECTION WITH THE SOFTWARE OR THE USE OR OTHERDEALINGS IN THE SOFTWARE.

Stringencoders License

Copyright 2005, 2006, 2007

NickGalbreath -- nickg [at] modp [dot] com

All rights reserved.

Redistribution and use in source and binary forms, with or without modification, arepermitted provided that the following conditions are met:

Redistributions of source code must retain the above copyright notice, this list ofconditions and the following disclaimer.

Redistributions in binary form must reproduce the above copyright notice, this list ofconditions and the following disclaimer in the documentation and/or other materialsprovided with the distribution.

Neither the name of the modp.com nor the names of its contributors may be used toendorse or promote products derived from this software without specific prior writtenpermission.

THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS ANDCONTRIBUTORS "AS IS" AND ANY EXPRESS OR IMPLIED WARRANTIES,INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OFMERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE AREDISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORSBE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, ORCONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENTOF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; ORBUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OFLIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDINGNEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THISSOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

This is the standard "new" BSD license:

http://www.opensource.org/licenses/bsd-license.php

zlib License

Copyright (C) 1995-2013 Jean-loup Gailly and Mark Adler

Architecting the Future of Big Data

Hortonworks Inc. Page 83

This software is provided 'as-is', without any express or implied warranty. In no event will theauthors be held liable for any damages arising from the use of this software.

Permission is granted to anyone to use this software for any purpose, including commercialapplications, and to alter it and redistribute it freely, subject to the following restrictions:1. The origin of this software must not be misrepresented; you must not claim that you

wrote the original software. If you use this software in a product, an acknowledgmentin the product documentation would be appreciated but is not required.

2. Altered source versions must be plainly marked as such, and must not bemisrepresented as being the original software.

3. This notice may not be removed or altered from any source distribution.

Jean-loup Gailly Mark Adler

[email protected] [email protected]

Apache License, Version 2.0

The following notice is included in compliance with the Apache License, Version 2.0 and isapplicable to all software licensed under the Apache License, Version 2.0.

Apache License

Version 2.0, January 2004

http://www.apache.org/licenses/

TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION1. Definitions.

"License" shall mean the terms and conditions for use, reproduction, and distributionas defined by Sections 1 through 9 of this document.

"Licensor" shall mean the copyright owner or entity authorized by the copyright ownerthat is granting the License.

"Legal Entity" shall mean the union of the acting entity and all other entities thatcontrol, are controlled by, or are under common control with that entity. For thepurposes of this definition, "control" means (i) the power, direct or indirect, to causethe direction or management of such entity, whether by contract or otherwise, or (ii)ownership of fifty percent (50%) or more of the outstanding shares, or (iii) beneficialownership of such entity.

"You" (or "Your") shall mean an individual or Legal Entity exercising permissionsgranted by this License.

Architecting the Future of Big Data

Hortonworks Inc. Page 84

"Source" form shall mean the preferred form for makingmodifications, including butnot limited to software source code, documentation source, and configuration files.

"Object" form shall mean any form resulting from mechanical transformation ortranslation of a Source form, including but not limited to compiled object code,generated documentation, and conversions to other media types.

"Work" shall mean the work of authorship, whether in Source or Object form, madeavailable under the License, as indicated by a copyright notice that is included in orattached to the work (an example is provided in the Appendix below).

"DerivativeWorks" shall mean any work, whether in Source or Object form, that isbased on (or derived from) theWork and for which the editorial revisions, annotations,elaborations, or other modifications represent, as a whole, an original work ofauthorship. For the purposes of this License, Derivative Works shall not include worksthat remain separable from, or merely link (or bind by name) to the interfaces of, theWork and Derivative Works thereof.

"Contribution" shall mean any work of authorship, including the original version of theWork and anymodifications or additions to that Work or Derivative Works thereof, thatis intentionally submitted to Licensor for inclusion in theWork by the copyright owneror by an individual or Legal Entity authorized to submit on behalf of the copyrightowner. For the purposes of this definition, "submitted" means any form of electronic,verbal, or written communication sent to the Licensor or its representatives, includingbut not limited to communication on electronic mailing lists, source code controlsystems, and issue tracking systems that are managed by, or on behalf of, theLicensor for the purpose of discussing and improving the Work, but excludingcommunication that is conspicuously marked or otherwise designated in writing by thecopyright owner as "Not a Contribution."

"Contributor" shall mean Licensor and any individual or Legal Entity on behalf of whomaContribution has been received by Licensor and subsequently incorporated withintheWork.

2. Grant of Copyright License. Subject to the terms and conditions of this License, eachContributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge,royalty-free, irrevocable copyright license to reproduce, prepare Derivative Works of,publicly display, publicly perform, sublicense, and distribute the Work and suchDerivative Works in Source or Object form.

3. Grant of Patent License. Subject to the terms and conditions of this License, eachContributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge,royalty-free, irrevocable (except as stated in this section) patent license to make, havemade, use, offer to sell, sell, import, and otherwise transfer the Work, where suchlicense applies only to those patent claims licensable by such Contributor that arenecessarily infringed by their Contribution(s) alone or by combination of theirContribution(s) with the Work to which such Contribution(s) was submitted. If Youinstitute patent litigation against any entity (including a cross-claim or counterclaim in alawsuit) alleging that the Work or a Contribution incorporated within theWork

Architecting the Future of Big Data

Hortonworks Inc. Page 85

constitutes direct or contributory patent infringement, then any patent licenses grantedto You under this License for that Work shall terminate as of the date such litigation isfiled.

4. Redistribution. You may reproduce and distribute copies of theWork or DerivativeWorks thereof in anymedium, with or without modifications, and in Source or Objectform, provided that You meet the following conditions:

(a) You must give any other recipients of theWork or Derivative Works a copy ofthis License; and

(b) You must cause any modified files to carry prominent notices stating that Youchanged the files; and

(c) You must retain, in the Source form of any Derivative Works that You distribute,all copyright, patent, trademark, and attribution notices from the Source form oftheWork, excluding those notices that do not pertain to any part of theDerivative Works; and

(d) If theWork includes a "NOTICE" text file as part of its distribution, then anyDerivative Works that You distribute must include a readable copy of theattribution notices contained within such NOTICE file, excluding those noticesthat do not pertain to any part of the Derivative Works, in at least one of thefollowing places: within a NOTICE text file distributed as part of the DerivativeWorks; within the Source form or documentation, if provided along with theDerivative Works; or, within a display generated by the Derivative Works, if andwherever such third-party notices normally appear. The contents of theNOTICE file are for informational purposes only and do not modify the License.You may add Your own attribution notices within Derivative Works that Youdistribute, alongside or as an addendum to the NOTICE text from theWork,provided that such additional attribution notices cannot be construed asmodifying the License.

You may add Your own copyright statement to Your modifications and may provideadditional or different license terms and conditions for use, reproduction, ordistribution of Your modifications, or for any such Derivative Works as a whole,provided Your use, reproduction, and distribution of the Work otherwise complies withthe conditions stated in this License.

5. Submission of Contributions. Unless You explicitly state otherwise, any Contributionintentionally submitted for inclusion in theWork by You to the Licensor shall be underthe terms and conditions of this License, without any additional terms or conditions.Notwithstanding the above, nothing herein shall supersede or modify the terms of anyseparate license agreement you may have executed with Licensor regarding suchContributions.

6. Trademarks. This License does not grant permission to use the trade names,trademarks, service marks, or product names of the Licensor, except as required forreasonable and customary use in describing the origin of theWork and reproducingthe content of the NOTICE file.

Architecting the Future of Big Data

Hortonworks Inc. Page 86

7. Disclaimer of Warranty. Unless required by applicable law or agreed to in writing,Licensor provides the Work (and eachContributor provides its Contributions) on an"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, eitherexpress or implied, including, without limitation, any warranties or conditions of TITLE,NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A PARTICULARPURPOSE. You are solely responsible for determining the appropriateness of usingor redistributing the Work and assume any risks associated with Your exercise ofpermissions under this License.

8. Limitation of Liability. In no event and under no legal theory, whether in tort (includingnegligence), contract, or otherwise, unless required by applicable law (such asdeliberate and grossly negligent acts) or agreed to in writing, shall any Contributor beliable to You for damages, including any direct, indirect, special, incidental, orconsequential damages of any character arising as a result of this License or out of theuse or inability to use theWork (including but not limited to damages for loss ofgoodwill, work stoppage, computer failure or malfunction, or any and all othercommercial damages or losses), even if such Contributor has been advised of thepossibility of such damages.

9. AcceptingWarranty or Additional Liability. While redistributing theWork or DerivativeWorks thereof, You may choose to offer, and charge a fee for, acceptance of support,warranty, indemnity, or other liability obligations and/or rights consistent with thisLicense. However, in accepting such obligations, You may act only on Your ownbehalf and on Your sole responsibility, not on behalf of any other Contributor, and onlyif You agree to indemnify, defend, and hold each Contributor harmless for any liabilityincurred by, or claims asserted against, such Contributor by reason of your acceptingany such warranty or additional liability.

END OF TERMS AND CONDITIONS

APPENDIX: How to apply the Apache License to your work.

To apply the Apache License to your work, attach the following boilerplate notice, withthe fields enclosed by brackets "[]" replaced with your own identifying information.(Don't include the brackets!) The text should be enclosed in the appropriate commentsyntax for the file format. We also recommend that a file or class name and descriptionof purpose be included on the same "printed page" as the copyright notice for easieridentification within third-party archives.

Copyright [yyyy] [name of copyright owner]

Licensed under the Apache License, Version 2.0 (the "License"); you may notuse this file except in compliance with the License. You may obtain a copy of theLicense at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributedunder the License is distributed on an "AS IS" BASIS, WITHOUTWARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.

Architecting the Future of Big Data

Hortonworks Inc. Page 87

See the License for the specific language governing permissions and limitationsunder the License.

This product includes software that is licensed under the Apache License, Version 2.0 (listedbelow):

Apache HiveCopyright © 2008-2015 The Apache Software Foundation

Apache SparkCopyright © 2014 The Apache Software FoundationApache ThriftCopyright © 2006-2010 The Apache Software FoundationApache ZooKeeperCopyright © 2010 The Apache Software Foundation

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this fileexcept in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under theLicense is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONSOF ANY KIND, either express or implied. See the License for the specific languagegoverning permissions and limitations under the License.

Architecting the Future of Big Data