Showing posts with label contract. Show all posts
Showing posts with label contract. Show all posts

Thursday, February 04, 2010

Why contracts are more important than designs

Following on from my last post on why IT evolution is a bad thing I'll go a stage further and say that far too much time is spent on designing the internals of elements of services and far too little on their externals. Some approaches indeed claim that working on those sorts of contracts is exactly what you shouldn't do as its much better for the contract to just be "what you do now" rather than having something fixed.

To my mind that view point is just like the fake-Agile people who don't document because they can't be arsed rather than because they've developed hugely high quality elements that are self-documenting. Its basically saying that everyone has to wait until the system is operable before you can say what it does. This is the equivalent of changing the requirements to fit the implementation.

Now I'm not saying that requirements don't change, and I'm not advocating waterfall, what I'm saying is that as a proportion of time allocated in an SOA programme the majority of the specification and design time should be focused on the contracts and interactions between services and the minority of time focused around the design of how those services meet those contracts. There are several reasons for this
  1. Others rely on the contracts, not the design. The cost of getting these wrong is exponential based on the number of providers. With the contracts in place and correct then people can develop independently which significantly speeds up delivery times and decreases risk
  2. Testing is based around the contracts not the design. The contract is the formal specification, its what the design has to meet and its this that should be used for all forms of testing
  3. The design can change but still stay within the contract - this was the point of the last post
The reality however is that IT concentrates far too much on the design and coding of internals and far too little on ensuring the external interfaces are at least correct for a given period of time. Contracts can evolve, and I use the term deliberately, but most often older contracts will still be supported as people migrate to newer versions. This means that the contracts can have a significantly longer lifespan than the designs.

As people rush into design and deliberately choose approaches that require them to do as little as possible to formally separate areas and enable concurrent development and contractual guarentees they are just creating problems for themselves that professionals should avoid.

Contracts matter, designs are temporary.

Technorati Tags: ,

Is IT evolution a bad thing?

One of the tenants of IT is that enabling evolution, i.e. the small incremental change of existing systems, is a good thing at that approaches which enable this are a good thing. You see it all the time when people talk about Agile and code quality and clearly there are positive benefits to these elements.

SOA is often talked about as helping this evolutionary approach as services are easier to change. But is the reality that actually IT is hindered by this myth of evolution? Should we reject evolution and instead take up arms with the Intelligent design mob?

I say yes, and what made me think that was reading from Richard Dawkins in The Greatest Show on Earth: The Evidence for Evolution where he points out that quite simply evolution is rubbish at creating decently engineered solutions
When we look at animals from the outside, we are overwhelmingly impressed by the elgant illusion of design. A browsing giraffe, a soaring albatross, a diving swift, a swooping falcon, a leafy sea dragon invisible amoung the seaweed [....] - the illusion of design makes so much intuitive sense that it becomes a positive critical effort to put critical thinking into gear and overcome the seductions of naive intuition. That's when we look at animals from the outside. When we look inside the impression is opposite. Admittedly, an impresion of elegant design is conveyed by simplified diagrams in textbooks, neatly laid out and colour-coded like and engineer's blueprint. But the reality that hits you when you see an animal opened up on a dissecting table is very different [....] a haphazard mess that we actually see when we open a real chest.


This matches my experience of IT. The interfaces are clean and sensible. The design docs look okay but the code is a complete mess and the more you prod the design the more gaps you find between it and reality.

The point is that actually we shouldn't sell SOA from the perspective of evolution of the INSIDE at all we should sell it as an intelligent design approach based on the outside of the service. Its interfaces and its contracts. By claiming internal elements as benefits we are actually undermining the whole benefits that SOA can actually deliver.

In otherwords the point of SOA is that the internals are always going to be a mess and we are always going to reach a point where going back to the whiteboard is a better option than the rubbish internal wiring that we currently have. This mentallity would make us concentrate much more on the quality of our interfaces and contracts and much less on technical myths for evolution and dynamism which inevitably lead into a pit of broken promises and dreams.

So I'm calling it. Any IT approach that claims it enables evolution of the internals in a dynamic and incremental way is fundamentally hokum and snake oil. All of these approaches will fail to deliver the long term benefits and will create the evolutionary mess we see in the engineering disaster which is the human eye. Only by starting from a perspective of outward clarity and design and relegating internal behaviour to the position of a temporary implementation will be start to create IT estates that genuinely demonstrate some intelligent design in IT.


Technorati Tags: ,


PS. I'd like to claim some sort of award for claiming Richard Dawkins supports Intelligent Design

Monday, July 21, 2008

Thinking about service levels

One piece that I've banged on about consistently for the last few years has been the importance of SLAs. SLAs that define not simply things like security and reliability (technical elements) but which also define the business contract and costs. These contracts must always work two ways. The consumer has a contract of things they must do to invoke the capability and have certain expectations of what they want. The service must have a contract of what it commits to do and what are the costs (time, financial, etc) for doing so.

In part this for me is about making pieces like network latency visible to developers and architects but more so making the actual business costs and value visible.

As an example take Paris, that city of surly waiters, stunningly dressed women and the complete decampment of people in August (lots of restaurants and shops are closed in August). Now every tourist wants to "go up the Eiffel tower" but if we look at the contracts here then there are other options that you could take from a "business" perspective.

What do you want to really do? The answer is get a birds eye view of Paris and see all of its wonders laid out before you. Now what are the wonders? The Louvre, Notre Dame, The eiffel tower, arc de triomphe. Now I'd argue that there is a service out there which costs less, in both time and money, which provides a better view, because the Eiffel tower is in it and which therefore represents the best business choice.

So the choice is basically queueing at the Eiffel tower

Or viewing at the Tour Montparnasse


The point here is that its a contract that enables this visibility and choice to the developer/architect/business rather than it being something implicit in the behaviour of a service. If you don't know what the contract is and most importantly what its price is (remembering that time is a cost after all) then you can't make a rational decision. Note here that I'm not talking about UDDI style dynamic discovery and binding I'm talking about making business decisions on a contractual basis.

Another example from Paris (if you haven't been to Paris BTW, stay in the 6th when you do, its "real" Paris in terms of being what you imagine). Now people catch trains and often you want to have a meal or a snack before you travel. Each of these places offers a varying type of contract which broadly covers the following
  • Quality of food
  • Quality of surroundings
  • Speed of service
  • Cost
Now lots of train stations have fast food joints, sit down light bite places and even sometimes something where you can get a reasonable meal. The point as a consumer is that my contract is defined by when my train leaves as well as my budget, this impacts what service I will use. If I know about the services before hand however I can plan to arrive at an earlier point if there is something I want to go to.

Now in Paris at Gare de Lyon there is "Le Train Bleu" this is just about the best damned railway restaurant in existence. They have a set of different menus including a "short" (by French standards) one for those who are about to travel.

The point here is that the contract for what could seem to be the same service (food at a railway station) can vary hugely from fast food through to a gourmet experience in an amazingly decorated french restaurant the contract isn't simply about price, or time, or quality, or trust its about a combination of those things and then weighing up which makes the most sense from a business perspective based on the current demands and ambitions of the business.

Exposing network effects to developers and architects is a tiny part of the problem, the real problem is exposing them to the business costs of their decisions and that requires a different degree of formalism and planning.

Once contracts are formalised it becomes possible to automate but the first key is being able to understand the contracts. There is still no WS-SLA or WS-Contract and the current push appears to be away from considering these elements and back to considering only technical aspects of service.

So think about the full contract not just on the immediately obvious this means not following the tourists in wasting 1/2 a day queueing for the Eiffel tower, but going to Montparnasse and having time to take in the cemetery and Jardin de Luxembourg and finishing up at Saint Sulpice having an espresso watching the beautiful people walk by. Now that is quality of service.

Technorati Tags: ,

Wednesday, June 18, 2008

Google App Engine - Quota limits

With my first Google App Engine application I deliberately decided to pick on something that is CPU heavy but which when threading is supported will go horizontally to help see how "wide" the cloud is.

Or maybe not.

Checking my results on usage I've come to the conclusion that a thread gets killed at around 9.5 seconds, so its on the deep calculation ones or a very big image. However this isn't the only quota that exists.


So to be clear this is the smallest mandelbrot set, its 200x200 and at this stage its a max of 16 iterations. This still means a potential for a lot of calculations but it really does indicate that Google are aiming at the data access end of the market rather than a commodity platform for doing heavy calculations (sort of a small problem competitor to IBM's Blue Gene.

In terms of what this means well on my MacBook Pro a "time python mandelbrot.py" request (which has to include starting python (about 0.2 seconds appears about right) which creates the "default" image takes the following

real 0m3.721s
user 0m3.436s
sys 0m0.097s

So with the python engine piece take off we have around 3.3 of user + sys and according to the logs this same request via the web is 9.4 times over quota. This means that the request quota is about 0.35 seconds of raw grunt but it will let you hang around in the data access for up to that 9.5 seconds.


Now right now Google are playing very nice with me given that pretty much every request passes or smashes the quota. Some go a LONG way past


The next zoom in then failed (at a recorded 31.7 times over quota! but around 9.5 seconds total time).

So fair play on Google for letting me abuse their infrastructure and it will be interesting to see what they do around CPU heavy tasks especially those that could go horizontally. I could see a real market for companies who effectively want 10,000 CPUs for a short period of time to calculate a forecast or similar where right now the cost/benefit analysis does stack up for a full hardware purchase but buying a whole load of Gigacycles (horizontally scaled) would be a great fit.



Technorati Tags: ,

Monday, April 28, 2008

SLAs and reporting - whose truth to believe?

One of the core concepts in SOA is the idea that a service should have a Service Level that it agrees to meet (the SLA), this might be technical in terms of up times, response times, amount of information to be handled. In more sophisticated services it could be more business oriented in terms of cost to serve, order to ship time, conversion rate, stock levels etc. These agreements can apply both ways, as in the service commits to respond within 20ms but the consumer can't send more than 20 messages a second.

SLAs are a guarantee that something will be done and there should be penalties in place when a violation occurs. In theory this should be a simple case of measurement, but all too often this is something that is overlooked. People define the SLAs but forget the old adage on KPIs that if you allow someone to measure their own KPIs they will be always be successful.

I've been looking for a simple demonstration of this problem for a while, most of them are specific to a given business so don't work generically. But thanks to the folks in Redmond I've now got a great example, it may or may not be there fault but its a good example of measurement problems.

I run Windows XP under Parallels on a MacBook Pro. Now I just do basic office work so I've given it a C: drive with 15GB. This should be a decent amount for the basic office files (Outlook files are stored on a dedicated share). The trouble is I've run out of space....
So what files are taking up the space? Well a quick "WinDirStat" on the C drive (after doing the same exercise with Windows properties) came up with the stat that the files on the drive take up around 5.4GB. This leaves me with around 10GB unaccounted for.
Here is an example of a producer/consumer SLA. The producer (Windows XP) has committed to provide me with around 15GB of storage for my office files. It then reports that I have violated the client SLA for the service and it will now perform like a dog. Equally my information is that the performance of the producer has severely degraded and I am unable to add more files despite being significantly below the agreed limit.

In this case it is the producer who measures the KPIs and even though the independent measure suggests it to be incorrect it is not possible to challenge the producers statement.

What this says is that when looking at KPIs and SLAs for services you need to think about independent measurement being part of the basic requirement. This implies that measurement is done at the service boundary by a 3rd party which must track these SLAs over time. Otherwise you'll just end up with Windows saying "disk space full please free up some space" and then telling you that you don't have enough files available to make a dent in the space.

The final piece therefore is around arbitration. It is quite possible in the setup that I'm using that something is going screwy around Windows that isn't the fault of the producer or the consumer its due to a 3rd party (e.g. the VM) once you get into this stage you need to think about arbitration. The challenge here is that the consumer is still due compensation from the producer, but the producer may have a counter claim against the 3rd party. This is the final piece about SLAs in a professional SOA environment. Its important to think "back to back" in your SLAs otherwise they won't actually mean anything. If a service commits to responding in 20ms but none of the things it relies upon will make any such guarantee then its just corporate optimism (at best) or fraud (at worst).

SLA management is a core part of the shift of SOA away from technology and into the business domain. With all the WS-* arguments going around I'm still stunned that there is no WS-Contract or WS-SLA because its that which would really separate WS-* from other technical choices.

SLAs are about
  1. Defining the terms
  2. Defining the penalities
  3. Measuring the operation
  4. Arbitrating the violation
  5. Spreading the risk down the chain
Its the independent measurement that helps make all the rest of it honest.

(Oh and if anyone has any idea about the 10GB I'd love to know where its gone!)

Technorati Tags: ,

Wednesday, April 16, 2008

Formalism over hacking - contracts, compilers and why engineering matters

At Uni I learnt Ada, in my first job I did Eiffel (and C) and in my second I did Ada (and C) as a result of which a bunch of people sent me a link to this article on how Ada delivered the sort of success that C programmers can only dream about. Udi Dahan then had a post over on InfoQ about how a "traditional" approach failed and required a much more engineering solution and some smart contracts.

Now I'm not going to say that Udi's project would have worked better first time in Ada (it wouldn't) but what it does highlight is something that Ada really helped people to focus on..... what is the range of acceptable results. Udi's project suffered because the expected range of data (tens) and the actual (hundreds of thousands) were so massively different.

The solution that Udi's team came up with appears pretty masterful and a great example of scaling a solution by looking at how the big clumps of data are actually assembled. By looking at the ranges and the manner in which they were created they were able to create a really great engineering solution.

Over on the Ada side the article highlights how Ada really upfronts the cost of development by making people really consider what is possible, what is probable and what shouldn't be allowed. By having these contracts it means that everything is explicit and this produces a greater focus on quality at those early stages of the project.

My experience with Ada was that it not only helped in ensuring this quality but that it really helped in ensuring the quality of the average developers on the team. Bluntly point an average developer in Ada was a productive member of the project, an average developer in C was a complete liability.

With all the talk of dynamic languages, scripting and late validation it is worth considering that when it really has to work it is always best to take an engineering approach and look at the contracts and ranges of information and to restrict the system to the correct handling of ranges rather than assuming that the system will be able to correctly handle everything.

As SOA becomes a standard way of operation and the issues of networked computing increase a focus on quality and contracts will become ever more important.


Technorati Tags: ,

Tuesday, February 19, 2008

English and American or why formalism is better than English

Lets be clear, I'm a fan of formalism. The best pure programming language I've ever used was Eiffel, the best programming language I've ever used in a large team for delivering the best code out of average developers and minimising issues was Ada. At Uni I had to learn Z notation and formal proofs and I've even used some of that stuff in certain highly reliable systems I've worked on down the years.

Simply put I think that having a formal specification that can be measured and adhered to is good engineering practice and that it reduces the impact of muppetry and misunderstanding. So that is my cards on the table, I like formalism because its consistently worked in large and complex projects where the skills profile has been mixed and the number of companies involved has been large.

Two countries separated by a common language - George Bernard Shaw

I travel a lot so it always entertains me when people argue that formal specifications aren't as effective as "simple" English specifications. Putting apart the UML/OCL challenge this has some basic problems in that what English do you mean? "British" English? American English? Scots English? Welsh English? Irish English? Australian English?


When an Australian says that you "need to wear thongs at the beach because its hot on the sand" they mean that you should wear things like this on your feet. A Brit will think they mean something like this(link) which is more than a little different.


When a Brit talks about a "bungalow" the Americans look on with same sort of confusion that the Brit has when an American says that someone "fell on their fanny" in polite society. When Americans refer to someone as "liberal" meaning it as an insult the Brit hears it as at worst a statement that the person is a bit wet.

The problem is that language is open to huge interpretation based on locality, even within single countries there can be different meanings. Describing someone as a "Baggie" in West Bromwich means they are a loyal and reliable person, ten miles down the road it means they are criminal scum who can't be trusted.

This isn't about pronunciation (for instance to all Bartenders in America, its highly unlikely that a Brit is ordering a "bear") its about the actual semantics of specific individual words that can completely change the meaning and outcome of a sentence. One example that was used at Uni to explain what a bugger English was is the following phrase

"The pen flew across the page"

Now it makes perfect sense right? It means someone is writing quickly, but could it mean something else?

Starting with the first noun what definitions can we have?

Pen - A writing implement, a place to keep animals, a female swan

then to the last noun

Page - a piece of paper, a small boy

then the tricky middle section

"flew across" - went in the air over, went fast

So which is more logical from the original statement?

The female swan went in the air over the small boy

or

The writing implement went fast over the piece of paper

Now we know its the last one right? But legally if it came to the argument could you explain why it could never be the former? This is the issue with "plain" English contracts, they are open to interpretation and this interpretation leads to delays and errors, formal specifications help to minimise this scope and thus help to improve quality. There is no such thing as "plain" English when it can be used to create nonsense poems like "Jabberwocky" and result in lawyers arguing over the interpretation of specific words in contracts and laws.

Put it this way, if someone wandered up to a bloke in an club or bar in the UK and said "Excuse me mate, can I bum one of your fags?" they'd be fine, in certain areas of the US they'd have a life expectancy of less than 30 seconds.

Technorati Tags: ,

Tuesday, January 22, 2008

Invocation v Intention

Just a quick note here on Invocation v Intention. In the SOA Methodology there is the "Why" bit that talks about why services (and actors) communicate.

I thought it was worth briefly explaining the difference between invocation and intention in these communications.

The invocation is the actual act itself, it is one service initiating the communication and asking for a specific capability to be delivered. Simplistically the invocation could be seen as an event reception by a service, a message being sent from one service to another or a standard procedure call. The point is that the invocation is about the capability on the destination service.

The intention is why the consumer made the call in the first place and why the producer wanted them to. What was their reason for doing it? Now the uber simplistic view is "to cause the capability to be delivered" but that is just "to get to the other side" to chickens. It doesn't really explain why an invocation was made.

So lets take why sales talk to finance. Nominally the invocation is "Report Sale" or some such. This can have some nice pre-conditions and post-conditions around what that invocation will do but it doesn't explain why Sales every actually calls the capability.

The reason Sales makes the invocation can vary. If they are bonused then the intention of the invocation is "to get paid my bonus". If they are measured on the reporting time then its "to comply with policy". Understanding these differences helps you understand what IT needs to do. Where it needs to smooth the way (policy) and where crappy IT will be overcome by the end users (bonus). This isn't to say you can deliberately have crappy IT if the sales guys are bonused but it does mean that while the users will gripe about the issue they will be personally driven to overcome it.

The other bit is on Finance. What is the intention of finance when they receive that report, why do they want them? Partly it is to measure the sales team and partly to make the finance departments life easier by having accurate numbers. The point here is this.

If your sales numbers suck because of a poor IT solution then you can look at the intentions and see if they can be changed to get around the IT, e.g. if the sales team aren't bonused can you bonus them in a way that is cheaper than replacing the IT system but means the reporting accuracy goes up. If you can't then its quite clear that the IT project to improve the sales reporting to finance isn't something that Sales will be very interested in (after all its a Finance policy they are complying with), this means that the people who will pay for the project will probably be Finance. If the current Sales system does a good job for the sales people but gives crappy reports for Finance then there isn't a motivation for Sales to focus on that area.

This is why its important to understand both the Invocation/Capability and the Intention when modelling services and their interactions.



Technorati Tags: ,

Saturday, September 15, 2007

Why contracts help migration

Now there are some people who think that validation is a bad idea, but recently I've seen a few cases where contracts have really come to the fore again. One of these is around migration. The scenario is pretty simple, you want to start moving to a new solution area but you don't want to do it in one go, you want to clearly split the area into two pieces, the consumer bit and the producer bit. Most likely its the producer that is going to be replaced but you don't want to like the two parts tightly together so they have to upgrade at the same time and the same pace.

So what do you do? Well first off you define a formal contract between the two areas, one that can be fulfilled by the current implementation and which can also be supported by the new service when it arrives. Effectively you've now split up the two worlds and placed an agreed boundary between them which can be enforced, tested against and measured against. You are now in the situation where the new implementation has to meet the contract and where the consumer can now rely on that contract independently of the implementation.

When looking at splitting up areas, or migrating applications and software the enforcement of clear contracts is a very simple way to start introducing business centric contracts and to provide a controlled way in which upgrades can be handled for both producers and consumers. Not using contracts means there isn't something to test and validate against and thus the new implementation is aiming more at an abstract goal than a concrete reality and the consumers have nothing to rely on as they plan their own future improvements.

Contracts therefore increase the flexibility of complex systems explicitly because they enforce the boundaries.

Technorati Tags: ,

Tuesday, August 28, 2007

Hotel Baths and the importance of client contracts

What is it with American Hotels in their competition to create large rooms with tiny baths. I've been in loads of Hotels and practically none of them in the US have a bath that is worth of the name "large sink" would describe most of them.

This is a cracking example of why a simple service definition isn't enough to ensure a decent quality of service. After all a bath is something that holds water, in which you can sit down, so these enlarged sinks match the basic criteria. They hold water, and if I get in one I get wet.

Now if when I specific a hotel room I could say

Bath.length > 1.2m
Bath.depth > 50cm

Thus giving a detail precondition on what I consider most important in a hotel room. I could also specify the details on the gym, to make sure its not just two crap machines in a tiny room.

Effectively here I talking about the importance of preconditions from a client perspective when selecting and invoking a service. A more business analogy would be when thinking about acceptable response times for a service, where if they don't respond in time you don't care about the result. Its about how the client can place restrictions and contracts onto the service.

The normal way to do this is via parameters on a method, but a more effective way would be via a contract applied to the call which could be managed as a policy independently of the invocation, i.e. changing acceptable response times.

Thus invocations should be about a negotiation between client and service to determine if the required conditions and QoS can be met, if they can't then the invocation is invalid, something the client should have as much say in as the service.

Seriously though, what is with US hotels? The beds are huge, the baths are tiny.



Technorati Tags: ,

Monday, July 30, 2007

Why Service SLA depends on consumers as much as producers (Sandisk Ultra II v Extreme III)

One of the bits that is often over looked in SOA is the importance of the consumer to an overall SLA for a service (after all the SLA is there to give the consumer assurance). I've been looking for a good example of what I mean and today I found it. I have a Canon EOS 350D, the old entry level digital camera, which is perfect for what I need. I've got 3 different brands of memory cards, all of which publish their performance SLAs. The two that interest me here though are the SanDisk Ultra II and SanDisk Extreme III. The later claims a minimum write speed twice that of the former. However if you lob it into my camera, set it on manual with a 1/4000 shutter speed and then hold down the button for 1 minute you get exactly the same number of shots out. This indicates that the issue isn't with the speed of the memory cards to determine overall performance but the speed of the camera in writing to the memory.

So here we have two services which have identical (one could say "uniform" :) ) interfaces but offer two different SLAs to consumers, these SLAs however place requirements on their consumers in order to be able to deliver the SLAs as advertised.

The point here is that a service cannot offer an SLA that reaches outside of its bounds without also placing direct requirements onto its consumers. Thus a true SLA is not simply something a service describes but is something that is negotiated between the producer and consumer for a given interaction.

Published SLAs are just the tip of the iceberg when it comes to delivering truly effective business SLAs that can be relied upon.


Technorati Tags: ,