Wednesday, 12 June 2013

Installing CouchDB - Get Training with Real time POC Project

Installing CouchDB

 CouchDB allows you to write a client side application that talks directly to the Couch without the need for a server side middle layer, significantly reducing development time. With CouchDB, you can easily handle demand by adding more replication nodes with ease. CouchDB allows you to replicate the database to your client and with filters you could even replicate that specific user’s data.
Having the database stored locally means your client side application can run with almost no latency. CouchDB will handle the replication to the cloud for you. Your users could access their invoices on their mobile phone and make changes with no noticeable latency, all whilst being offline. When a connection is present and usable, CouchDB will automatically replicate those changes to your cloud CouchDB.
CouchDB is a database designed to run on the internet of today for today’s desktop-like applications and the connected devices through which we access the internet.

Step 1 

The easiest way to get CouchDB up and running on your system is to head to CouchOne and download a CouchDB distribution for your OS — OSX in my case. Download the zip, extract it and drop CouchDBX in my applications folder (instructions for other OS’s on CouchOne).
Finally, open CouchDBX.

Step 2 – Welcome to Futon

After CouchDB has started, you should see the Futon control panel in the CouchDBX application. In case you can’t, you can access Futon via your browser. Looking at the log, CouchDBX tells us CouchDB was started at http://127.0.0.1:5984/ (may be different on your system). Open a browser and go tohttp://127.0.0.1:5984/_utils/ and you should see Futon.

CouchDB jQuery Plugin

Futon is actually using a jQuery plugin to interact with CouchDB. You can view that plugin athttp://127.0.0.1:5984/_utils/script/jquery.couch.js (bear in mind your port may be different). This gives you a great example of interacting with CouchDB.

Step 3 – Users in CouchDB

CouchDB, by default, is completely open, giving every user admin rights to the instance and all its databases. This is great for development but obviously bad for production. Let’s go ahead and setup an admin. In the bottom right, you will see “Welcome to Admin Party! Everyone is admin! Fix this”.
Go ahead and click fix this and give yourself a username and password. This creates an admin account and gives anonymous users access to read and write operations on all the databases, but no configuration privileges.
Users in CouchDB can be a little confusing to grasp initially, specially if you’re used to creating a single user for your entire application and then managing users yourself within a users table (not the MySQL users table). In CouchDB, it would be unwise to create a single super user and have that user do all the read/write, because if your app is client-side then this super user’s credentials will be in plain sight in your JavaScript source code.
CouchDB has user creation and authentication baked in. You can create users with the jQuery plugin using$.couch.signup(). These essentially become the users of your system. Users are just JSON documents like everything else so you can store any additional attributes you wish like email for example. You can then use groups within CouchDB to control what documents each user has write access to. For example, you can create a database for that user to which they can write to and then add them to a group with read access to the other databases as required.

Step 4 – Creating a Product Document

Now let’s create our first document using Futon through the following steps:
  1. Open the mycouchshop database.
  2. Click “New Document”.
  3. Click “Add Field” to begin adding data to the JSON document. Notice how an ID is pre-filled out for you, I would highly advise not changing this. Add key “name” with the value of “Nettuts CouchDB Tutorial One”.
  4. Make sure you click the tick next to each attribute to save it.
  5. Click “Save Document”.

Step 5 – Updating a Document

CouchDB is an append only database — new updates are appended to the database and do not overwrite the old version. Each new update to a JSON document with a pre-existing ID will add a new revision. This is what the automatically inserted revision key signifies. Follow the steps below to see this in action:
  • Viewing the contents of the mycouchshop database, click the only record visible.
  • Add another attribute with the key “type” and the value “product”.
  • Hit “Save Document”.

Step 6 – Creating a Document Using cURL

I’ve already mentioned that CouchDB uses a RESTful interface and the eagle eyed reader would have noticed Futon using this via the console in Firebug. In case you didn’t, let’s prove this by inserting a document using cURL via the Terminal.
First, let’s create a JSON document with the below contents and save it to the desktop calling the fileperson.json.
  1. {  
  2.     "forename": "Gavin",  
  3.     "surname":  "Cooper",  
  4.     "type":     "person"  
  5. }  
Next, open the terminal and execute cd ~/Desktop/ putting you in the correct directory and then perform the insert with curl -X POST http://127.0.0.1:5984/mycouchshop/ -d @person.json -H "Content-Type: application/json". CouchDB should have returned a JSON document similar to the one below.
  1. {"ok":true,"id":"c6e2f3d7f8d0c91ce7938e9c0800131c","rev":"1-abadd48a09c270047658dbc38dc8a892"}  
This is the ID and revision number of the inserted document. CouchDB follows the RESTful convention and thus:
  • POST – creates a new record
  • GET – reads records
  • PUT – updates a record
  • DELETE – deletes a record

Step 7 – Viewing All Documents

We can further verify our insert by viewing all the documents in our mycouchshop database by executingcurl -X GET http://127.0.0.1:5984/mycouchshop/_all_docs.

Step 8 – Creating a Simple Map Function

Viewing all documents is fairly useless in practical terms. What would be more ideal is to view all product documents. Follow the steps below to achieve this:
  • Within Futon, click on the view drop down and select “Temporary View”.
  • This is the map reduce editor within Futon. Copy the code below into the map function.
    1. function (doc) {  
    2.     if (doc.type === "product" && doc.name) {  
    3.         emit(doc.name, doc);  
    4.     }  
    5. }  
  • Click run and you should see the single product we added previously.
  • Go ahead and make this view permanent by saving it.
After creating this simple map function, we can now request this view and see its contents over HTTP using the following command curl -X GET http://127.0.0.1:5984/mycouchshop/_design/products/_view/products.
A small thing to notice is how we get the document’s ID and revision by default.

Step 9 – Performing a Reduce

To perform a useful reduce, let’s add another product to our database and add a price attribute with the value of 1.75 to our first product.
  1. {  
  2.     "name":     "My Product",  
  3.     "price":    2.99,  
  4.     "type":     "product"  
  5. }  
For our new view, we will include a reduce as well as a map. First, we need to map defined as below.
  1. function (doc) {  
  2.     if (doc.type === "product" && doc.price) {  
  3.         emit(doc.id, doc.price);  
  4.     }  
  5. }  
The above map function simply checks to see if the inputted document is a product and that it has a price. If these conditions have been met, the products price is emitted. The reduce function is below.
  1. function (keys, prices) {  
  2.     return sum(prices);  
  3. }  
The above function takes the prices and returns the sum using one of CouchDB’s built in reduce functions. Make sure you check the reduce option in the top right of the results table as you may otherwise be unable to see the results of the reduce. You may need to do a hard-refresh on the page to view the reduce option.

Get Hands-on Training @ BigDataTraining.IN
Contact us:

#67,2nd Floor, 1st Main Road, Gandhi Nagar, Adyar, Chennai- 600020




Friday, 31 May 2013

Apache Zookeeper Training @ BigDataTraining.IN

Apache ZooKeeper is an effort to develop and maintain an open-source server which enables highly reliable distributed coordination.
ZooKeeper is a centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services. All of these kinds of services are used in some form or another by distributed applications. Each time they are implemented there is a lot of work that goes into fixing the bugs and race conditions that are inevitable. Because of the difficulty of implementing these kinds of services, applications initially usually skimp on them ,which make them brittle in the presence of change and difficult to manage. Even when done correctly, different implementations of these services lead to management complexity when the applications are deployed. 

http://www.hadooptrainingchennai.in/courses/

http://www.hadooptrainingchennai.in/hadoop-training/
 


ZooKeeper is a high-performance coordination service for distributed applications. It exposes common services - such as naming, configuration management, synchronization, and group services - in a simple interface so you don't have to write them from scratch. You can use it off-the-shelf to implement consensus, group management, leader election, and presence protocols. And you can build on it for your own, specific needs.

ZooKeeper: A Distributed Coordination Service for Distributed Applications

ZooKeeper is a distributed, open-source coordination service for distributed applications. It exposes a simple set of primitives that distributed applications can build upon to implement higher level services for synchronization, configuration maintenance, and groups and naming. It is designed to be easy to program to, and uses a data model styled after the familiar directory tree structure of file systems. It runs in Java and has bindings for both Java and C.

Design Goals

ZooKeeper is simple. ZooKeeper allows distributed processes to coordinate with each other through a shared hierarchal namespace which is organized similarly to a standard file system. The name space consists of data registers - called znodes, in ZooKeeper parlance - and these are similar to files and directories. Unlike a typical file system, which is designed for storage, ZooKeeper data is kept in-memory, which means ZooKeeper can achieve high throughput and low latency numbers.
The ZooKeeper implementation puts a premium on high performance, highly available, strictly ordered access. The performance aspects of ZooKeeper means it can be used in large, distributed systems. The reliability aspects keep it from being a single point of failure. The strict ordering means that sophisticated synchronization primitives can be implemented at the client.
ZooKeeper is replicated. Like the distributed processes it coordinates, ZooKeeper itself is intended to be replicated over a sets of hosts called an ensemble.
ZooKeeper is ordered. ZooKeeper stamps each update with a number that reflects the order of all ZooKeeper transactions. Subsequent operations can use the order to implement higher-level abstractions, such as synchronization primitives.
ZooKeeper is fast. It is especially fast in "read-dominant" workloads. ZooKeeper applications run on thousands of machines, and it performs best where reads are more common than writes, at ratios of around 10:1.

Data model and the hierarchical namespace

The name space provided by ZooKeeper is much like that of a standard file system. A name is a sequence of path elements separated by a slash (/). Every node in ZooKeeper's name space is identified by a path.

The ZooKeeper Data Model

ZooKeeper has a hierarchal name space, much like a distributed file system. The only difference is that each node in the namespace can have data associated with it as well as children. It is like having a file system that allows a file to also be a directory. Paths to nodes are always expressed as canonical, absolute, slash-separated paths; there are no relative reference. Any unicode character can be used in a path subject to the following constraints:
  • The null character (\u0000) cannot be part of a path name. (This causes problems with the C binding.)
  • The following characters can't be used because they don't display well, or render in confusing ways: \u0001 - \u001F and \u007F - \u009F.
  • The following characters are not allowed: \ud800 - uF8FF, \uFFF0 - uFFFF, \uXFFFE - \uXFFFF (where X is a digit 1 - E), \uF0000 - \uFFFFF.
  • The "." character can be used as part of another name, but "." and ".." cannot alone be used to indicate a node along a path, because ZooKeeper doesn't use relative paths. The following would be invalid: "/a/b/./c" or "/a/b/../c".
  • The token "zookeeper" is reserved.
      

    Getting Started with ZooKeeper

    Standalone Operation

    Setting up a ZooKeeper server in standalone mode is straightforward. The server is contained in a single JAR file, so installation consists of creating a configuration.
    Once you've downloaded a stable ZooKeeper release unpack it and cd to the root
    To start ZooKeeper you need a configuration file. Here is a sample, create it in conf/zoo.cfg:
    tickTime=2000
    dataDir=/var/lib/zookeeper
    clientPort=2181
    
    This file can be called anything, but for the sake of this discussion call it conf/zoo.cfg. Change the value of dataDir to specify an existing (empty to start with) directory. Here are the meanings for each of the fields:
    tickTime
    the basic time unit in milliseconds used by ZooKeeper. It is used to do heartbeats and the minimum session timeout will be twice the tickTime.
    dataDir
    the location to store the in-memory database snapshots and, unless specified otherwise, the transaction log of updates to the database.
    clientPort
    the port to listen for client connections
    Now that you created the configuration file, you can start ZooKeeper:
    bin/zkServer.sh start
    ZooKeeper logs messages using log4j -- more detail available in the Logging section of the Programmer's Guide. You will see log messages coming to the console (default) and/or a log file depending on the log4j configuration.
    The steps outlined here run ZooKeeper in standalone mode. There is no replication, so if ZooKeeper process fails, the service will go down. This is fine for most development situations, but to run ZooKeeper in replicated mode,

    Managing ZooKeeper Storage

    For long running production systems ZooKeeper storage must be managed externally (dataDir and logs).


    In machine learning and pattern recognition, a feature is an individual measurable heuristic property of a phenomenon being observed. Choosing discriminating and independent features is key to any pattern recognition algorithm being successful in classification. Features are usually numeric, but structural features such as strings and graphs are used in syntactic pattern recognition.
    The set of features of a given data instance is often grouped into a feature vector. The reason for doing this is that the vector can be treated mathematically. For example, many algorithms compute a score for classifying an instance into a particular category by linearly combining a feature vector with a vector of weights, using a linear predictor function.
    The concept of "feature" is essentially the same as the concept of explanatory variable used in statistical techniques such as linear regression.
     
    BigDataTraining.IN Technology focus towards project development and professional training in Big Data and Hadoop Technologies. We served the students with our academic projects.
    Machine Learning Training Chennai with POC Projects !
     
    We were able to prove our worth in the following areas.

    Hadoop
    Big Data
    Big Data Analytics
    Big Data & Hadoop Development solutions
    Advanced Hadoop EcoSystems Tools
    MongoDB
    Apache Cassandra
    HBase – Developer & Admin
    Sentiment Analysis
    Prediction Engine
    Recommendation Engine
    Mahout
    CouchDB
    HBase
    CouchBase
    Prediction Engine
    Cloud Computing
    VMware
    Xen
    KVM
    Amazon EC2
    Eucalyptus
    Open Stack
    Android
    IOS-IPHONE
    Mobile Computing

    We assist more number of people with our student projects and provide exposure and support to the students with our Technical Architects every year. Lot of scholars from various colleges and universities are benefitted and hence, we still receive referrals from engineering colleges all over India.

    Learn Big Data from Big Data Solutions Architects! Hadoop Training Chennai with 
    Hands-On Practical Approach ! Reach us to Enroll! 100% Placements
     
    Key Features -
    Cloud Server Access
    Training = Enterprise Scale
    Advanced Technology Coverage + PoC Project Work
    24/7 Technical Support

    http://www.bigdatatraining.in/machine-learning-training/


    http://www.bigdatatraining.in/hadoop-development/training-schedule/

    Mail:
    info@bigdatatraining.in
    Call:
    +91 9789968765
    044 - 42645495

    Visit Us:
    #67, 2nd Floor, Gandhi Nagar 1st Main Road, Adyar, Chennai - 20
    [Opp to Adyar Lifestyle Super Market]