Saturday, June 25, 2016

Data Types, Sessions and Other Basic Usages of TensorFlow


TensorFlow is a computational framework that allows users to represent computation as a graph. Nodes in this computational graph represents an operation to be performed, which is called “op” in TensorFlow context. An op can take zero or more Tensors and produce zero or more Tensors as it’s output. Tensor can be referred to as a multi-dimensional array. 

In this post I’ll be focusing on Computational graphs, Sessions, Tensors, and Variables.

Typical TensorFlow graph has three main phases, which are Construction phase, Graph Assemble phase and Execution phase. By default TensorFlow provides a computational graph where users can assign ops that needs to be executed without assembling a new graph. 



Operations and Tensors

A Tensor represents the data flow through ops, in other words a value consumed or produced by an op. Tensors does not hold the value of that operation, instead it handles as a symbolic handle to an output of an op. In this perspective, Tensor plays two roles; it can be passed as an input to an op which allows to build the graph and it’s data flow, Secondly it allows to perform computations on values being passed. In below, matrix_1, matrix_2 and product are Tensors.

matrix_1 = tf.constant([[3., 3.]])
matrix_2 = tf.constant([[2.], [2.]])
product = tf.matmul(matrix_1, matrix_2)

Like mentioned previously an operation or op is a node in TensorFlow graph that can both consume or produce zero or more Tensors. Ops performs computations on Tensors. For example,

product = tf.matmul(matrix_1, matrix_2)

Creates an operation of type “matmul” and takes matrix_1 and matrix_2 tensors as inputs to produce product tensor as output.

Graphs

A graph in TensorFlow can be referred as a collection of Ops that linked by Tensors and can be start with ops that doesn’t take any input and pass their output to other ops that performs computations. In below example we have added three ops including two constants and one op to TensorFlow default graph. To launch and get the results we must execute the graph in a TensorFlow session (more on Sessions later).


matrix_1 = tf.constant([[3., 3.]])
matrix_2 = tf.constant([[2.], [2.]])
product = tf.matmul(matrix_1, matrix_2)

sess = tf.Session()

result = sess.run(product)

Sessions

A session contains all operations and tensors of a computational model. It does execute ops and evaluates tensors. The default session can be launched without giving any graph arguments when creating the session. If it’s required to run different graphs in different sessions, the we have to provide which graph should be executed in each session we are creating.

TensorFlow session represents a connection to it’s underline native libraries. So it holds all resources allocated until we close the sessions. Or else we could user session as a context manages.

matrix_1 = tf.constant([[3., 3.]])
matrix_2 = tf.constant([[2.], [2.]])
product = tf.matmul(matrix_1, matrix_2)

sess = tf.Session()

result = sess.run(product)

print(result)
sess.close()

To execute above example with session as a context manager,

matrix_1 = tf.constant([[3., 3.]])
matrix_2 = tf.constant([[2.], [2.]])
product = tf.matmul(matrix_1, matrix_2)

with tf.Session() as sess:
result = sess.run(product)

print(result)

Variables

Variables keeps state across TensorFlow graph executioin and can be considered as in-memory buffers of Tensors. Variables should be explicitly initialized and can be stored in the hard disk to use them later on.

To create a variable we must pass a Tensor as it’s input parameter to the constructor.

x = tf.Variable([1.0, 2.0])
a = tf.constant([3.0, 3.0])

# create a variable that will be initialize to scalar value 0.
state = tf.Variable(0, name="counter")

# variables must be initialized by running an 'init' op after having
# launch the graph. first add the init op to the graph.
init_op = tf.initialize_all_variables()

In another post, we will be dive deep into data flow graphs for better understanding of computation executions.

For source code: https://github.com/isurusiri/tensorflowbasics

Setting Up TensorFlow Environment for CPU Based Usages


Deep learning has done wonders work in both Natural Language Processing and Image Processing in last few years. Even though some of these algorithms has been existed for decades, recently performance have been improved and ultimately proved its capabilities by the improvements in hardware infrastructure and accumulated huge data volumes. 

This improvements has lead development and research communities to introduce frameworks consist of deep learning algorithms. Caffe, CNTK, DeepLearning4J, Torch, Theano and TensorFlow are some of the very famous deep learning libraries at the moment.

First things first, I am not a deep learning expert. Luckily I had the opportunity work very closely with some deep learning focused researches soon after I graduated. 

I have been working with Convolutional Neural Networks and Autoencoders in last few months and I was evaluating few of above mentioned libraries, but TensorFlow seemed to like the Chuck Norris of deep learning libraries. In this post I will give the instructions to setup TensorFlow for CPU base usages and run the “Hello world!”.



In here I assumes the environment is Ubuntu and have installed Python 2.7 or 3.0+.

First install the PIP package manager.

$ sudo apt-get install python-pip python-dev

Set and environment variable for the Tensorflow binary release (here I have added 64bit Python 2.7 version).

$ export TF_BINARY_URL=https://storage.googleapis.com/tensorflow/linux/cpu/tensorflow-0.9.0-cp27-none-linux_x86_64.whl

Just install the TensorFlow now.

$ sudo pip install --upgrade $TF_BINARY_URL

To test the TensorFlow, open terminal and execute,

$ python
...
>>> import tensorflow as tf
>>> hello = tf.constant('Hello, TensorFlow!')
>>> sess = tf.Session()
>>> print(sess.run(hello))
Hello, TensorFlow!
>>>

That’s it, everything should work perfectly. Note that you might have to remove any installations if you have installed TensorFlow before.

In next post I’ll introduce basic usages of TensorFlow.

For source code: https://github.com/isurusiri/tensorflowbasics


Sunday, May 1, 2016

The True OLAP

OLAP or simply OnlineAnalytical Processing is the process that allows users to conduct multidimensional interaction analysis operations on real-time business data. This includes performing operations such as drilling, aggregating, pivoting, and slicing multidimensional data.

It is necessary to create topic specific data cubes in advance to support above operations, thereby users can inspect data in tables or graphs and conduct real time pivoting and drilling. Having said that, let’s consider whether OLAP is enough for catering analysis and forecasting in real world.

Figure 1: Information Systems can be divided into transnational (OLTP) and analytical (OLAP). In general OLTP provide source data to data warehouse, whereas OLAP helps to analyze it.


A mature company may have somewhat large data accumulated about its operations. These data can be used to make certain guesses about the business they are engage in. For example, a vehicle importer may guess what kind of people are tending to buy what kind of vehicles. These guesses are just the basis for forecast. Then the company could utilize accumulated data to evaluate above guesses. When a guess evaluated to be true they can be used in forecast and when it is false they will be re-guessed.

Above process could be referred as evaluation process, whose purpose is to justify conclusions with evidence find in historical data. In business analysis process, a query like the first n customers who has purchased a vehicle from vehicle types contributed for half of the sales volume of the company in the year the x is ubiquitous and required some form of a computation or querying, with intermediate steps.

The requirement of building the data cube in advance, and a limited set of actions available to perform against data cube are limiting the analysis process in a situation like above. The data model needs to have capabilities of reconstructing the cube or temporarily build cubes to cater diverse analysis demands. Furthermore, most of the OLAP products are more famous for their rich user interface; only a few has powerful online analytical capabilities.

Figure 2: In OLAP database there is aggregated, historical data, stored in multi-dimensional schemas (usually star schema).

So what kind of an online analytical tool could fulfill the evaluation process? Theoretically, steps for evaluation can be considered as computation regarding data. This computation can be defined by user and can decide next computation actions to be taken based of the intermediate results without having a defined model beforehand. Additionally this computation should support performing actions on huge amount of data instead of simple numeric computations. At this point of view, SQL is somewhat fulfilling this requirement, but considering its own computational capability, still it has its own limits on solving problems similar to above mentioned problem. This leaves us in a position to think about the limitations of SQL and builds a way through it to make a new generation of computational system for evaluation process, namely, the real OLAP.

I’m still completely a novice in Business Intelligence. J

Saturday, April 23, 2016

What Makes Business Intelligence Better?

I’m completely a Novice in Business Intelligence J

Business Intelligence is the process of giving businesses (technically to people who steers the business) insights about what they have been doing in their operations. This is done by processing data gathered during business transactions. This data is mostly in unstructured manner. This unstructured data available in internal and external data sources are transformed into reportable forms such as graphs as tables. The end result of business intelligence process is a set of reports and/or a dashboard that allows non technical business operators to perform ah-hoc queries on their data.

When we hear the term Business Intelligence, more precisely Intelligence, we tend to have an impression of performing huge amount of analytical tasks and serve a set of predictions. Bazinga! This is not the case in business intelligence at all. Yes, Corporations need numbers, facts about how they are performing in the process of achieving KPIs, but most importantly the data should make sense about operations to non technical users. This is where a business intelligence tool gets success, because not every user is going to be a node js geek who is waiting for npm to finish up his work.

Ideally the tool should support to build data visualizations in the matter of minutes, or maybe seconds without requiring technical report development knowledge. For example having functionalities like drag-drop might be quite interesting. In terms of reporting, the tool should require a minimal number of steps/clicks to build a report.

Ability to introduce new fields based on calculations using other existing fields/columns, filtering data based on predefined or auto modeled parameters, exploring data from different angles, and suggesting data types, schema and hierarchies automatically are also important in self service aspects of the tool.

Business intelligence also needs to perform predictive analytics. The tool should be capable of analyzing historical and current data to make predictions about future events. It should have functions to more clearly communicate a huge amount of complex data.
Furthermore it is important that tool could handle data sets in different sizes. The volume of data being consumed and number of queries being run simultaneously greatly affects the performance of the tool.



Ultimately business intelligence tool should support to interrogate data and draw conclusions that help to perform job better. Having a trending software stack doesn’t make the product a better one. A sustainable set of technologies that facilitates the product to get job done is what allows it to stand ahead in the game. 

Wednesday, February 24, 2016

Eastern and Western Narratives of the Philosophy of Language

Language and linguistics are two interesting research areas that I’d like to focus. These interests were stimulated by the work I have done in natural language processing, like unambiguously interpreting instructions from user manuals. After resigning from my previous work place I had some time to restard blogging the things I’m interested. Here I summed few facts that I have collected time by time and I’m hoping to write more about language and linguistics often.

Language is one of the most expressive forms of communicating knowledge with rest of the world; a master of this expressive form could share his wisdom with others. Philosophy of language can be denoted as a systematic study/inquiry of the origins of language, the nature of meaning, the usage and cognition of language and the relationship between language and reality.

The curious fate of humankind.
Philosophy of language has its roots in final centuries BC and early centuries AD. VedicAge in Indian history made contribution to the philosophy of language in early stages of this quest, to be precise philosophical schools of Nyaya and Mimamsa gave rise to linguistic philosophy. For example Sanskrit GrammaticalTradition can be considered and one such contribution.

Ancient Greece covered most of work in philosophy of language from the Western tradition. Plato, Aristotle and the Stoics made contributions in that era.

Plato believed that names of things are based on the nature where each smallest structural unit that distinguishes meaning represents basic ideas or sentiments. Aristotle considered that the way a subject is modified or described in a sentence is established through an abstraction of the similarities between various individual things. The Stoic philosophers separated five parts of speech: nouns, verbs, appellatives, conjunctions and articles. Their contributions are mostly made to the analysis of grammar.

In late 19th Century language start playing a major role in Western philosophy. Publication like “Cours delinguistique generale” made an impact to that contribution. 

Saturday, November 28, 2015

I Took the Red Pill; Setting Up Development Environment to Contribute Wildfly (1)

I have been working with JBoss EAP and Wildfly  for almost 2 years and the experiences is quite interesting. Inside JBoss EAP it uses Wildfly as its core JEE server. What I like most in Wildfly  are its easy configuration and management interface, modular design and indeed fast startup.


There were times I had to extensively investigate class loading mechanisms and dependency management with related to Wildfly, which gave me sort of a high confidence about its runtime environment. Its just crazy and made me to dig deep. Here I’m explaining steps required to setup your development environment to start Wildfly contribution.

Prerequisites,
  • Java 8
  • Mavan 3
  • GitHub account 

Setting up a personal repository to work,

1. Fork Wildfly repository into your personal account.
2. Clone your newly forked repository to local workspace.  

$ cd wildfly

3. Add a remote reference to upstream, for pulling future updates .

$ git remote add upstream git://github.com/wildfly/wildfly.git

4. Disable merge commits to your master.

$ git config branch.master.mergeoptions –ff-only

Build the Source Code,

Building WildFly 9 requires Java 8 or newer, make sure you have JAVA_HOME set to point to the JDK8 installation. Build uses Maven 3.

1. Run build.sh script.

$ ./build.sh

but,
Basically all you do is run

$ mvn clean install

or if your don’t want to wait for default testsuites you can do

$ mvn install -Dmaven.test.skip=true

Hakuna matata!


Launch Wildfly,

Built Wildfly zip file can be find in [repository_location]\build\target\

1. Navigate to above mentioned directory.

$ cd build\target\wildfly-10.0.0.CR5-SNAPSHOT\bin

Wildfly supports comes with two operation modes; standalone and domain.
To run Wildfly in standalone mode,

2. Execute standalone.sh script.

$ ./standalone.sh



From another post I’ll explain how to load Wildfly source to IntelliJ Idea and start debugging.



Monday, November 16, 2015

Machine Learning for Fluid Rendering in Video Games


Data-driven Fluid Simulations using Regression Forests from SoHyeon Jeong on Vimeo.

Water is one of elements that 3D animators put a lot of effort to get the perfection. It is really difficult to simulate dynamic movements of millions of particles and it requires huge amount of computational resources, specially for real time video rendering in 3D games.

However a group of researchers from Swiss Federal University of Technology and Disney Research has published a paper that shines some light on above problem with the aid of Machine Learning.

Rather than calculating the movements of each particle or reducing the particle count, this method could predict the dynamic movements of particles accurately. Presented algorithm needs to be trained with a random set of videos which have animated fluid particles with accurate calculations. Furthermore, the algorithm doesn't calculate movements in real-time, rather it predicts according to prior knowledge.

According to researches, this algorithm can render fluid animations 3 times faster than existing methods and can animate nearly 2 million particles in real-time.