It’s amazing how some concepts take off like gangbusters in a short duration of time. Big Data is one such concept, that creeps into our conversations because of all the market noise. There is definitely merit to the fundamental premise behind Big Data for most businesses; create better end-user experience, make intelligent business decisions, reduce intellectual waste and monetize on new opportunities or opportunities that did not present itself before. Thus the demand for Data Scientists, application developers, statisticians, mathematicians, etc. -- note these are mostly on the development and analytic side of the house. What’s amazing is large databases have been there for the longest time, in many cases, even the data that are targets now for Big Data applications were also available for the longest time. What has evolved rapidly are the applications tools that facilitate optimized manipulation of massive data sets and flexible interfaces to diverse databases -- example Hadoop.
You may have heard that the digital universe is in petabytes, global IP traffic is in 100s of exabytes. These are mind bogglingly large metrics. Big data analytics can play a crucial role in making datasets in this space usable – by improving operational efficiency to customer experience to prediction accuracy. While Cisco is the global leader in networking -- Did you know that 85% of estimated 500 exabyte global IP traffic in 2012 will pass through Cisco devices ? – the company also builds an innovative family of unified computing products. This enables the company to provide a complete infrastructure solution including compute, storage, connectivity and unified management for big data applications that reduce complexity, improves agility, and radically improves cost of ownership.
To meet a variety of big data platform demands (Hadoop, NoSQL Databases, Massively Parallel Processing Databases etc), Cisco offers a comprehensive solution stack: the Cisco UCS Common Platform Architecture (CPA) for Big Data includes compute, storage, connectivity and unified management. Unique to this architecture is the seamless data integration and management integration capabilities with enterprise application ecosystem including Oracle RDBMS/RAC, Microsoft SQL Server, SAP and others. See Figure 1.
The Cisco UCS CPA for Big Data is built using the following components:
- Cisco UCS 6200 Series Fabric Interconnects provides high speed, low latency connectivity for servers and centralized management for all connected devices with UCS Manager. Deployed in redundant pairs offers the full redundancy, performance (active-active), and exceptional scalability for large number of nodes typical in big data clusters. UCS Manger enables rapid and consistent server integration using service profile, ongoing system maintenance activities such as firmware update operations across the entire cluster as a single operation, advanced monitoring, and option to raise alarms and send notifications about the health of the entire cluster.
- Cisco UCS 2200 Series Fabric Extenders, act as remote line cards for Fabric Interconnects providing a highly scalable and extremely cost-effective connectivity for large number of nodes.
- Cisco UCS C240 M3 Rack-Mount Servers, 2-RU server designed for wide range of compute, IO and storage capacity demands. Powered by two Intel Xeon E5-2600 series processors and support up to 768 GB of main memory (typically 128GB or 256GB for big data applications) and up to 24 SFF disk drives in the performance optimized option or 12 LFF disk drives in the capacity optimized option. Also features Cisco UCS VNIC optimized for high bandwidth and low latency cluster connectivity with support for up to 256 virtual devices.
Read More »
Last night we uploaded version 6.1 of Cisco’s Tidal Enterprise Scheduler. I’m pretty excited to introduce the new functionality of this tool and there’s a lot. Particularly with Hadoop support and Amazon EC and S3 support as well. If you are unfamiliar with TES, the datasheet is here.
But when talking about big data, I thought, I’d start small. Like iPhone small. Existing Scheduler customers and the curious, can download the free Apple iPhone app to control jobs. Here’s the AppStore description and link
Cisco Enterprise Scheduler is the premiere job scheduling and process automation software that provides a single point of control and monitoring for business operations. Enterprise Scheduler for iOS now allows Scheduler administrators and users to monitor and control their operations directly on their mobile devices. Enterprise Scheduler for iOS was designed for the mobile user experience, but retains core features of the Enterprise Scheduler web client that users are familiar with including:
* Monitor and view jobs, connections, events, schedules, queues, logs and alerts.
* Control all aspects of jobs, including holding, rerunning, canceling, and overriding jobs.
* Powerful search and filtering for all Scheduler objects.
Last week we participated in the annual Hadoop Summit held in San Jose, CA. When we first met with Hortonworks about the Summit many months back they mentioned this year’s Hadoop Summit would be promoting Reference Architectures from many companies in the Hadoop Ecosystem. This was great to hear as we had previously presented results from a large round of testing on Network and Compute Considerations for Hadoop at Hadoop World 2011 last November and we were looking to do a second round of testing to take our original findings and test/develop a set of best practices around them including failure and connectivity options. Further the set of validation demystifies the one key Enterprise ask “Can we use the same architecture/component for Hadoop deployments?”. Since a lot of the value of Hadoop is seen once it is integrated into current enterprise data models the goal of the testing was to not only define a reference architecture, but to define a set of best practices so Hadoop can be integrated into current enterprise architectures.
Below are the results of this new testing effort presented at Hadoop Summit, 2012. Thanks to Hortonworks for their collaboration throughout the testing.
Expanding its Big Data portfolio, Cisco announced a fully integrated end-to-end hardware and software infrastructure for enterprise Hadoop deployments in partnership with Greenplum, a division of EMC, that delivers industry-leading performance, scalability, advanced management capabilities and enterprise-class service and support. This solution consists of Cisco UCS 6200 Series Fabric Interconnects, Cisco UCS C-Series rack mount servers and Greenplum MR. Greeplum MR is based on the MapR M5 distribution, a completely re-engineered implementation of the Apache Hadoop stack with 100 percent compatibility. Cisco UCS is the exclusive integrated platform for Greeplum MR that can significantly reduce time-to-value and the operating expenses associated with Hadoop implementations.
Hadoop implementations can present a number of challenges to enterprise environments, many of these arise from the dichotomy between the introduction of innovative new technology and the enterprise-class performance, reliability, and support demanded by mission-critical systems. The collaboration between Cisco and Greenplum is specifically designed to provide a solution to these challenges. The joint solution delivers radically simplified deployment and management, high availability, excellent performance, exceptional scalability, and world-class service and support from long-time collaborators Cisco and EMC.
This solution can also connect, across the same management plane, to other Cisco UCS deployments running enterprise applications, thereby radically simplifying data center management and connectivity.
The configuration starts in a single rack with the ability to extend into multiple racks.
For more information or deal inquiries, please email us at: firstname.lastname@example.org. A joint white paper is available at http://www.cisco.com/en/US/solutions/collateral/ns340/ns517/ns224/ns944/wp_greenplum.pdf.