Friday, February 24, 2006

Don't Ask, Do Tell

There's a great piece up at The Pragmatic Programmer about some very basic, but very important OOP concepts. Read the article for a complete explanation, but here's a very short summary...

  1. Don't ask an object for information about it's state and then make decisions based on that data. You shouldn't be making decisions about the state of an object and then changing that state yourself. This logic should be inside that particular object. Just tell it what you want done and then let it do that to itself.
  2. Don't tightly couple objects together by passing data around from one to another. Once you've returned some data from an object, any semantic meaning associated with it is lost. Basic OOP states that data should be stored along with it's state/methods; i.e. in its own object. Related to point 1.
  3. Don't use multiple methods to access/work upon a single invariant state variable. Multithreading will ensure that invariants can no longer be guaranteed invariable (unless you lock properly). Better to use something like a map() call to apply actions to data.
In other words, follow basic OOP fundamentals rigorously. Keep data and the methods/logic that act upon it together in one place. Don't scatter it all about.

Tuesday, February 21, 2006

S.P. Jain - Principles and Framework for Project Management

Five, 9 hour days of gruelling work later, I must admit that the course was totally worth it. The time literally flew by and I can say with a fair bit of confidence that my perspectives on the art and science of project management have shifted significantly.

I've learnt a lot and not just from the instructor (Prof. Ajay Parasrampuria). As he himself pointed out, with more than 200 man-years of experience in the room, there's not a whole lot he could have added. 'Course, he went right ahead and did just that! But more on that later...

There were around 20 participants from a wide range of backgrounds. We had people from USAID, GAIL and Emirates. Large corporations, small companies and one hopeful entrepreneur. People with more than 40 years of organisational management experience and those with very little.

And of course, me :-)

The most important and illuminating part of the entire 5 days was the sharing of past experiences by the participants. Some were more reserved than others and remained regrettably silent, but there was enough back and forth to keep everyone interested.

Now, I don't want to give the impression that we did all the work and Ajay just stood back and listened in :-). His knowledge of the PMBOK framework was outstanding, as was his ability to communicate with the rest of us. His ability to continually relate principles from the framework to his and others real life experiences was absolutely riveting.

Perhaps the most educational part of the course were the daily 'games' we participated in. Now I can't reveal the details since we were requested not to ruin the surprise for the next batch, but I hope I can say they were excellent learning experiences. For me, the moment of PMBOK epiphany came during one of them when I noticed just how easy managing the entire (rather complicated) project became once we abandoned our ad-hoc procedures and started to follow the process framework. Things started to slot into place almost immediately and it was shockingly easy to manage all that data and derive meaning from it.

The highlight of the course for me was the presentation on "Monitoring and Controlling IT Projects" by Mr. Tapan Bose, V.P. for Financial/Banking Projects at Satyam Computers. Lots of useful observations from the trenches and insights on what works and what doesn't. The main points I picked up was not to forget the human dimension (especially team building) and to treat the process as a part of the project, not as a template filling chore carried out under threat of audit.

Of course, the most depressing part of the entire exercise was the realisation by all concerned, of how little of what we'd learnt will actually be applied in real life. We might be all gung-ho about working according to the PMBOK framework in our next project, but without institutional support (what the PMBOK would call Enterprise Environmental Factors and Organisational Processes Assets :-), it's very difficult to be an agent of change. There has got to be buy in from top management, or it's all just an expensive waste of time.

Here's to hoping I can actually use what I've learnt!

Saturday, February 04, 2006

TurboGears - It's all in your head

Turbogears is absolutely, positively, definitely, unquestionably completely amazing!

I started out with very rusty python (I haven't coded in it for more than a year and they've managed to sneak in new language features in the meantime) and only a cursory knowledge of SQLObject, CherryPy and Kid (the three main components of the stack). Around 4 hours later, I have an almost working Todo List web application and a grin on my face that could eat a banana sideways.

The only reason it took me 4 hours is because of the ramp-up time associated with learning any new technology. I had to stop every few minutes and read the docs, scratch my head a little and then continue. And the only reason the application is 'almost working' is because I've decided to extend the capabilities significantly and so I'll be re-writing it.

But pause to consider what I've just accomplished. I've got the canonical 'Hello World' of web applications up and running (well, almost) in less than half a working day and I'm so charged up about it, I'm already thinking of version 2. A whole SDLC compressed into half a working day! And the size of the code! The entire controllers.py file is smaller than the Hibernate XML mappings I'd have had to write to get something similar in Java.

But the most important part of the experience is the pythonic characteristic of TurboGears.

You can hold it all in your head.

With most Java frameworks and ORMs, you're running to the docs every few minutes, every time. With TurboGears, now that I've done it once, I really don't have to go back to the docs at all; at least for the bits I've worked with. That's a real source of satisfaction and a major cause behind the speed boost.

Time to complete my next app of the same size and complexity? I'm guessing 45 minutes :-)

Thursday, February 02, 2006

Face Time and Free Stuff

I've blogged about the need to interact with people before here and here. Christopher Hawkins has written an excellent article that succinctly sums up what it's all about:

Face Time and Free Stuff

Tuesday, January 10, 2006

Ruby/Rails Redux

I tried. I really did.

If you've been following this blog, you'll know that I gave Ruby a look-see when the hype over Rails first started to build up. I didn't particularly like what I saw and being a bit of a Java weenie, I went back into its all encompassing, XML filled, embrace.

But I've absolutely had it with Java now. It's great for large complicated projects with teams of developers working together, but for a one man effort, cobbled together in what free time I have, Java is over-kill. It's official, I'm finally sick to death of the ginormous configuration files and over-engineered class hierarchies. Do I really need 10+ classes just to access the database? I hope not.

So I succumbed to the siren like call of the Rails fanboyz and decided to give Ruby on Rails a try... but I couldn't do it; I just couldn't. [cue weepy music]

The rational reason I'll trot out is that Ruby has terrible support for I18n/Unicode, but the real reason is that Ruby's syntax is just too much like Perl (aka 'line noise'). Enough to give me flashbacks about my time in the trenches writing that filth. They had to drag me out from under the bed and pull my thumb out of my mouth.

And don't get me started on the 'magical' ways of Rails. I hate config files as much as the next guy, but you can have too much of a good thing.

So what's the alternative?

Python.

Now, I don't need reminding that Python has it's warts, I've had to stare at them often enough. I dislike the explicit 'self' variable, the lack of access control modifiers and the almost-but-not-quite OOP schematics. But's it's leagues better than Ruby and it's got fairly decent I18n.

Pythons not been getting the buzz Ruby has as far as its frameworks go, but there are two that are certainly worth considering.

1. TurboGears
2. Django

Both promise Rails like speed of development and simplicity through a single integrated Web Application stack. TurboGears glues together various best of breed frameworks for it's needs while Django builds up everything from scratch.

Some people seem to prefer Django, but I haven't decided which way I'll go yet. This much however, is clear...

Ruby can ride the next train out of town and good riddance.

Tuesday, January 03, 2006

There are no uninteresting numbers


Jim Loy's theorem: There are no uninteresting numbers.

Proof: Assume that there are. Then there is a lowest uninteresting number. That would make that number very interesting... which is a contradiction.

Hmmm. Interesting... :-)

More at http://www.jimloy.com/math/math.htm


Monday, January 02, 2006

Friday, December 23, 2005

"My Golden Rule" -- Business 2.0 Magazine

My Golden Rule

49 business visionaries, collectively worth over $70 billion, state which single philosophy they swear by more than any other.

It's interesting, but as always take this stuff with a whole load of salt. Things are never as clear-cut as the blurb suggests. Who knows if it's these "rules" that made them successful or whether they just wish they had.

:-)

Are there no truly original ideas left?

Look it what I found. Openomy.

Let's see now...

  1. Online File System? Check.
  2. Tagging? Check.
  3. REST based? Check.
  4. 1 GB space free? Check.
  5. Transparent Integration coming? Check.
See this blog entry for an introduction to Openomy.

Sounds real similiar to this and this, doesn't it? :-)

Thursday, December 22, 2005

4 Rules for the Practical Entrepreneur -- Fragmented Markets

Ian Landsman has written an excellent article for the practical entrepreneur in this article. Definitely worth the time spent reading it.

Thursday, December 15, 2005

SLoPIS (sloppies): Slow Loading Pages of Infinite Size

I've sometimes wondered how I'd instantiate and maintain a call-back link between a client and a server on the Internet. That is, I want the server to update the client as and when various events occur without the client having to specifically poll for them. This is very easy to do over an intranet, but seems almost impossible over the Internet because most clients will be behind restrictive firewalls and proxies. Using an arbitrary protocol over a randomly chosen port is not going to work. So is using polling the only way to go?

One way out of this conundrum might be to use something I've christened SLoPIS (pronounced: sloppies) - for Slow LOading Pages of Infinite Size - in a moment of whimsy.

What your client will do is issue a GET command to a particular URL, which will be generating the actual call-backs. The server starts to send back a 'page' with any queued events or a ping every 5 seconds to keep the connection alive. To any entity in the middle, the page simply appears to be very large and very slow loading, but not otherwise unusual in any way.

So lets say I have a REST based file system or something and I want to be informed of any file changes on the server. I can open up a connection to http://some.filesystem.com/SLoPIS/filechanges which is, lets say, a Java Servlet on the server. As part of the request I send my authentication data and the server responds with the any queued events. e.g.

<html>
<body>
DELETION: /mnt/some/file
...

If there's nothing to report for awhile and the connection is in danger of timing out, it'll send a ping, like so:

...
PING
...

to continue the downloading of the 'page'.

Now this isn't a true call-back, because the connection has to be initiated by the client, but for most applications I can think of, this is not really such a problem. To receive, the client just keeps an ear to the sloppy and reacts to any events sent over it.

So no more issues with firewalls or proxies and no need to create a special hole in them to get your application to work! :-) I can see this being very valuable for AJAX based web-apps.

There can be issues with load and running out of ports from a client source IP if there are a lot of NAT'ed users coming in from the same IP, so it's not a perfect solution. However, it may be more appropriate than polling in many situations.

On Ennui (on-we)

"This is the basic state of the creative soul when work has no meaning and brings suffering. In this state, employees feel they are doing the same things day after day. They repeat the same tasks, fill out the same forms and talk to the same people. They work in an environment they cannot change. They are merely the executors of other’s projects. They believe their bodies are the extension of other people´s minds.

In this stage, employees are not supposed to be creative. Don’t think, just do it! They are told. They don’t feel their work is important because they feel replaceable. Finally, they are not satisfied with what they do. When someone spends half their life doing something they don’t enjoy, it impacts their soul. Indeed, it has a deep impact.

This is the stage of Sysiphus. Sysiphus was a character of ancient Greek mythology. He represents two states of the soul at work: suffering and lack of meaning. Albert Camus explained the fate of Sysiphus with these words:

“The gods had condemned Sisyphus to ceaselessly rolling a rock to the top of a mountain, whence the stone would fall back of its own weight. They had thought with some reason that there is no more dreadful punishment than futile and hopeless labor”.

But the soul tends to move. It is like a river that always finds a channel."

-- The Life Cycle of the Creative Soul

Wednesday, December 14, 2005

A tag based file system (with Bayesian Auto-tagging)

So having been foiled by OmniDrive in my desire to create an Internet based virtual drive, I've moved on other, perhaps greener and defintely less populated pastures.

If you'll remember, I spoke about adding tagging support to the virtual drive I was talking about. I said that we could "... add tagging support to the virtual drive. You can tag files and folders and view virtual 'tag' folders with links to those files. Mainstream OS's don't have a tagging mechanism for files, so we'll have to add meta-data through file names. e.g. end file names with a special character and the tags (i.e. myfile.txt#work,proposal,text) which will be stripped off before being saved to the virtual drive."

I won't be working on the 'Internet' part of the virtual drive, but I can certainly implement this idea. Why not create a FUSE based file system with tag based virtual folders? Use either folder names (/mnt/tagfs/here:are:some tags/myfile.txt) or file names (/home/user/myfile.txt:here:are:some tags) to add tags and then use virtual tag directories to navigate through those tags? You can use mv to change the tags associated with a file etc.

In addition, we can have a Bayesian Auto-tagger which learns which tags you've previously used for files of a certain type and then automatically tags them appropriately if no tags are supplied. The more you tag, the better the auto-tagger.

Who knows, I might decide to run with this one! :-)

The saga continues...

Relevant links:

Monday, December 12, 2005

Another one bites the dust

No sooner do I start thinking about a globally accessible, scaleable, encrypted virtual drive that someone announces their intention of releasing just such a product! :-(

Check out OmniDrive. Pretty much what I was aiming for. They've even got their eyes set on Google and have a very interesting blog entry on The Economics of Online Storage.

Well, back to the drawing board! :-)

Friday, December 09, 2005

The Fine Art of Programming

A fine and growing collection of online programming guides and such. A link worth saving and revisiting.

Thursday, December 08, 2005

A virtual drive in every pot

I've been mulling over the idea of a virtual drive on the Internet for some years now. Witness my ham-handed efforts at http://zfs.sourceforge.net for an early example. Well, I recently resurrected the idea of writing something like it now and it's interesting to see how my thoughts have evolved.

ZFS as I initially envisioned it was to be a network of automatically replicating file servers and the use case in my mind was a university file server. There would be a mapping of many users to a single (virtual) server, with the system (internally a cluster) having to scale to handle as close to an infinite number of users as possible.

Lately, I've been thinking more along the lines of writing a 'Net Drive' type application. A virtual disk I can mount from any machine connected to the Internet and treat as a local drive. Companies like xDrive and iDrive and MangoSoft already offer something of the sort. However, their offerings are targeted more towards the business user. I personally feel there's a massive untapped market of casual users who might be interested.

Imagine having 1GB of space available to you online and directly accessible via a virtual drive. Directly save documents, media files etc. to the virtual drive and access it from anywhere through another computer with the same drive mounted in or through a web interface. Share your password with several people and have them save in the same drive if you wish, give them a URL to the data on your drive or just share certain folders. Boom, you've just eliminated the need for a hard disk on our PC. Internet appliances, here we come!

This type of application would be ideal for someone like Google to create and I can see them stepping into this field sometime soon. It's a classic Google app. You need to scale almost infinitely, but that's easy because you can create slices of the virtual resource and limit the number of users accessing each slice. Want to support more users? Add more slices.

Take Gmail as an example. It's probably got hundreds of millions of users, but unlike the University use case, the users have a many to many (or from another perspective, a one to one) relationship with the system. That is, unlike university students, gmail users are not interested in checking other people's mail or accessing a common email account, or even sharing their email account. This makes it much easier to slice up the virtual space, assigning a limited number of users to each slice and scaling the slices. So Gmail is probably made up of thousands of individual computers, each supporting let's say 1000 users, fronted by an authentication cluster. When a user wants to log in, he goes to gmail.com, is authenticated and then redirected to the individual machine he shares with 999 other people. If you want to add another 1000 users, plug in another machine. You can keep scaling horizontally till infinity for all practical purposes. The authentication datastore will eventually become a bottle-neck, but you can support a enormous number of users before you hit that wall. *

If Google were making a virtual drive, they'd do something similiar. As new users signed up, they'd be assigned to different machines, upto a certain max cap. Just like gmail, users have a one to one relationship with their account. That is, they're only interested in the contents of their accounts and have no need to access anyone elses account or a common store. This makes it trivial to scale exponentially.

Now this is a great product to make and market, except for one small problem; there are already a whole bunch of people out there doing the same thing. So we need to differentiate ourselves from the pack.

One way to do that is to add tagging support to the virtual drive. You can tag files and folders and view virtual 'tag' folders with links to those files. Mainstream OS's don't have a tagging mechanism for files, so we'll have to add meta-data through file names. e.g. end file names with a special character and the tags (i.e. myfile.txt#work,proposal,text) which will be stripped off before being saved to the virtual drive. Users can also publicly 'share' tags.

Other features we can offer are:

  1. Fast file indexing and searching and maybe even mapping/linking files to each other based on content etc.
  2. Clients for hand-helds with disconnected operational ability
  3. Single-click integration with Flikr, Del.icio.us etc.
  4. Rsync based transfers
How are you going to pay for all this? Advertising. Have the virtual drive folder show text/banner ads and the website as well. Have premium accounts and dedicated machines for business users.

Who knows, I might work on this idea... or maybe not.

* Correction: It's possible to avoid turning the authentication store into a bottle-neck. One way to do this would be to have the store for a particular set of users reside on the machine assigned to them. So when you want to access the virtual drive abc.virtualdrive.com, you go to that URL and send in your username and password. If the authentication process running on that machine can't find the user, the login attempt fails.

This just leaves the DNS server as the bottleneck now :-)

Wednesday, December 07, 2005

PETA - People Eating Tasty Animals

If God didn't want us eating animals, why did He make them out of meat? :-D

A couple of friends and I were discussing the various methods used to slaughter animals (over lunch, when else) and we got to talking about the Halaal method (where a stroke through the neck severs every vein, artery and the wind-pipe, but leaves the spinal column intact) versus the Jhatka method (lit. jerk - where the animal is decapitated in one stroke). The debate revolved around which method caused the least pain to the animal, with most people automatically assuming that the Jhatka method was more painless.

I disagree.

In both methods, the animal eventually becomes unconscious due to a lack of blood going to the brain and the resultant drop in blood pressure. In the Jhatka method, since the animal is decapitated and the head is no longer attached to the body, we don't see the animal kick about and grunt as animals being slaughtered are wont to do. However, the sensation of 'pain' is interpreted by the brain and as long as that is active, the animal will suffer. The head being separated from the body doesn't make any difference.

Since in both methods, the blood flow is disrupted and in one we have the additional pain of the spine being cut through, the Jhatka method should logically be the more painful of the two, with the added disadvantage of not keeping the heart going for as long as possible to clear out as much of the pathogen carrying blood as possible.

It's just an illusion that Jhatka is more merciful, brought on by the stillness of the decapitated animal corpse. A final verdict awaits the time when it will be possible for us to measure and quantify 'pain' as a value...

Hungry kya? :-P

Tuesday, December 06, 2005

SQL Injection and XSS Attacks

Some topics everyone involved in web development must read at least once:

  1. SQL Injection Attacks by Example - It's a lot easier than you think. And yes, you customers will try it out some of the standard approaches out of idle curiosity if nothing else.
  2. Real World XSS - You'll be surprised at the sites which are vulnerable to attacks of this nature.
  3. More XSS
  4. And still more XSS

Mark Cuban - Success & Motivation

Blog Maverick - The Mark Cuban Blog

In Essence:

  1. Be driven
  2. Know the industry you're in, inside out.
  3. Use 1. and 2. to ensure you're ready when Lady Luck strikes.
  4. You only need to be lucky once...

Cuban comes across as an intense, driven, workaholic; kind of like my current (successfully entrepreneurial) employer :-D. It's a bit depressing to think that the only way to free yourself from the chains of a 9-5 life, is to handcuff yourself to a 00:00-23:59 one!

Friday, November 18, 2005

ColorMatch Redux

Thanks to this, I can now tell which shirt will go with which trouser.

Breakthrough!