Monday, December 12, 2022

The Future of Knowledge Commodification


A new way to package and sell information on a question-by-question basis

Ever smaller packages of commodifiable "knowledge"

The evolution of knowledge and its accessibility has come a long way. From the scale of books to daily papers, and then to on-demand search through the likes of Yahoo and Google, we have seen a steady increase in the availability of information. In the 2010s, the rise of 24-hour news and the ability to receive hourly updates directly on our phones has further enhanced our access to knowledge.

From Shared "Search" to Custom, On Demand Answers

Recently, traditional search has become even more refined, providing more accurate answers to specific questions. However, the next big development in the commodification of knowledge will come on December 1st, 2022 with the introduction of a new way to package and sell information on a question-by-question basis. This knowledge will be unique to the individual and the specific conversation and moment in time, and will not be shared or read by anyone else.

A "Futures Market" For Knowledge???

In the future, there will even be futures markets for knowledge, as business owners will want to forecast and lock in prices for the questions they will need answered in the next quarter. The advent of AI agents that can provide better knowledge at a lower cost will also be a game-changer in this industry. Overall, the commodification of knowledge on a sentence-by-sentence basis is a significant development that will greatly impact the way we access and utilize information. 

OpenAI Prioritizes Market Share Over Revenue


ChatGPT uses rate-limiting tactics to retain free users

Yesterday, it was suggested that OpenAI may need to raise their prices in order to address capacity issues. However, in a surprising turn of events, the company has instead reduced their prices, even for paid accounts, back down to zero. This move has infuriated paid users who are now being rate-limited in order to keep the service available for free customers.

It's easy to see the analogy to drug dealers giving new users a free taste, while telling the addicts that they're out of supply. But this strategy makes sense for OpenAI, as the lifetime value of those free customers far exceeds the few dollars they might be able to get from paid users over the next few days.

Saturday, December 10, 2022

The Rapidly Increasing Cost of Knowledge Itself, with ChatGPT


The world as we know it has changed... we just don't know it yet

Introduction

ChatGPT has changed the world forever by making knowledge a tradeable commodity, the cost of which is rising, while simultaneously revolutionizing and  likely leading to the collapsing in cost of software development over time

Knowledge: prices soaring,

Let's start with this: Last Week - Google.com was Free. 

Despite still being free Yesterday,
 I paid $4.96 for the privilege of being able to ask ChatGPT, OpenAI's astonishing  new Chatbot (chat.openai.com/chat) ~913 questions.  

Those same questions, same answers, i.e. the same "Knowledge", will cost me ~$6 today

What will it cost Tomorrow?  I suspect that literally not even OpenAI knows.

I'm sorry, but that is just BONKERS.

Software: costs likely to collapse

At the same time that the cost of accessing knowledge is clearly on the rise, the impact of AI technologies like ChatGPT on software development is likely to be a staggering reduction in cost over the coming months/years.


With the ability to automate research and design processes, as well as assist in implementation, projects that previously required multiple full-time employees can now be completed with fewer resources and in less time.

This article explores the implications of both the benefits, as this will likely lead to significant cost savings for businesses and organizations, as well as increased efficiency and productivity, it is also a huge shift to have to pay real dollars for the best knowledge of the day.

What is ChatGPT?

ChatGPT is a chatbot developed by OpenAI, a leading research institute in the field of artificial intelligence. It is a large language model trained using deep learning techniques, which allows it to generate human-like responses to natural language inputs. ChatGPT is able to answer a wide range of questions, providing accurate and detailed answers in real time.

The emergence of ChatGPT last week is significant because it is an example of a rapidly advancing AI technology that is changing the way knowledge is accessed and valued.

While slower versions of ChatGPT are available for free, the rapid increase in the cost of using the paid tier of ChatGPT to access knowledge is indicative of the broader trend of rising prices in the AI industry, and also seems like it has the potential to create huge economic and social inequalities move forward.

Potentially Alarming Implications

ChatGPT has followed this trajectory since Nov 30th; we are literally watching the price of knowledge rise from day-to-day, in real world dollars!

Date Range Price Response Time Responsiveness
Dec 1st - Dec 4th $0.0000 1 second Free! Works great.
Dec 5th - Dec 6th $0.0050 5-10 seconds Free Slow. Paid Okay.
Dec 7th - Dec 8th $0.0055 10-30 seconds Even Paid Slow.
Dec 9th - Dec 9th $0.0066 30-45 seconds Paid Really Slow.
Dec 10th... $0.00?? 60-75+ seconds ~1 request per min

In just a few days, the cost of accessing knowledge through ChatGPT has increased significantly.

2 days ago, the cost per question was roughly $0.004, but yesterday it had risen to $0.0054, and today it is more than $0.0066.

This obviously represents a significant increase over the previous cost of accessing knowledge through traditional search engines like Google, which was essentially free.

The Rapidly Changing Economic Landscape of AI

The Price of Knowledge in an AI World

This AI technology is quickly going to end up pricing folks that can't afford to pay for this knowledge out of the market.

Google was basically free.

Now suddenly from one day to the next, the price for knowledge literally has a tangible, real-world price.  And that price is currently still going up every day. I would not be shocked if this price were to even go as high as $0.05 per question in the short term.  At least that's my guess today. 

This guess is mainly based on the fact that I would happily pay that, with even a basic understanding of what this fucking thing is capable of ... having spent most of every waking minute talking to it - since learning of it's existence from my friend Edgar, at 2:44PM CST, Dec 3rd, 2022.  

Will AI Replace Developers?

One of the best quotes I've heard about AI to date is, was in a YouTube video from Anestasi in Tech:

AI isn't going to replace Developers.  
Developers who use AI are going to replace Developers who don't.

So to make that happen, "using AI", at even $.01 or $0.10/question will potentially cost $5-10 (or more) per day, per employee, just to acquire "the knowledge" that they will need in order to move at this new velocity.  

In addition to the new Velocity of Change that this AI allows for, and all of the changes to everything else in the world that are about to cascade from that change in velocity, there is also now potentially a completely new:

~$1-3K: "Knowledge Acquisition Cost", Per Employee, Per Year - just to keep up with the Joneses.

Adapting to the Challenges of AI Knowledge

It seems as though it is essentially an entirely different economic model, that the vast majority of the world is still completely unaware of let alone prepared for.  This is simultaneously both a little bit terrifying, and yet also feels a little bit like the internet must have in the early 90's. 


The difference is that in that case those folks had roughly 2-3 decades to respond.

This is (and will keep) moving SO ASTONISHINGLY FAST that, accounting for tech-inflation, we probably have more like 2-3 years at most, and possibly more like 2-3 months to adapt.

This thing is just so incomprehensively powerful.  The folks that learn how to, and who are privileged enough to be able to pay to use it, are going to just simply dominate those that don't. 

Society is Unlikely to Benefit Equally

The rapid rise in the cost of knowledge is not only alarming due to the rate of increase, but also because it is a significant departure from the previous model of mostly free to the world's access.  This is likely due, at least in part, to the high demand for this technology and the literal, logistical limitations of availability of the physical servers that support it.

As more and more people begin to use ChatGPT, the servers that host this technology are becoming strained, leading to an increase in the cost of access.  This trend is likely to continue for the foreseeable future, as demand for AI technologies continues to grow and the infrastructure struggles to keep up.

Unfortunately, the increasing cost of knowledge is likely already creating a divide between those who can afford to pay for access to this knowledge, and those who cannot. As the cost continues to rise, this divide is likely to become more pronounced, with potentially far-reaching implications for the economy and society as a whole.

The AI industry is already moving at an extraordinary pace, and this new knowledge acquisition cost only adds to the challenges and opportunities that it presents.

Imagine A $5K GPTesla!

On the positive side of things, as stated above, ChatGPT is likely to reduce the cost of software development over the coming months, and years, by one or more orders of magnitude.  It would be a little bit like if car's suddenly got 10x cheaper. 

Imagine waking up tomorrow, and finding that there's a new company called GPTesla with a car that has a 1000 mile range, charges from a wall socket in 30 seconds and costs $5K rather than $50k. 

Like, literally, from one day to the next.  As a software developer, that's the world I seem to find myself in today.  😲🤔😅😀

Conclusion

The rapidly increasing cost of knowledge is a worrying trend that has the potential to create a divide between those who can afford to access knowledge and those who cannot.

It is important that we are all aware of this. As the use of AI technology continues to grow, it will be interesting to see the impact on both the availability and cost of knowledge.

In the short term at least, I'm sorry, but this is going to be a metaphorical blood bath. 

I don't know how it is all going to play out, but **it is about to get real!

Monday, November 7, 2022

Characteristics of a Great Single Source of Truth

The Power and Magic of Knowledge Graphs comes from a few, essential elements that they must support.  By far the most import of these is queriability, but a close second is that there is only 1 of them - like literally, in the world ideally.

Some additional attributes that make for a good single source of Truth are:

  1. A great Single Source of Truth is, basically, a traditional RDBMs
  2. + 3 simple types of read-only, calculated fields
    1. Parent Lookup Values i.e. CustomerTaxRate = order.CustomTaxRate
    2. Child Aggregations - i.e. LineItemSubTotal = sum(lineItems.SubTotal)
    3. Simple, In-Row Calculations - i.e. Total = CustomerTaxRate * LineItemSubTotal
  3. The ability to export 100% of this data to a Json, XML, CSV or similar data storage format as a snapshot of the SSoT at any given moment in time.

Queriability

The main problem with language in general is that it is not "queriable" without an intelligent agent (either human or AI) to parse the "natural language text" into a model that can be queried with semantic precision.

Singleton

This makes the cloud a fantastic place for a Knowledge Graph, because it can act as a Single Source of Truth by construction, based on the fact that anyone in the globe has access to literally a single instance of the top level graph.

Normalized RDBMs

This makes it queriable, and normalized at it's heart, allowing for an N-dimensional structure.

3 simple types of Read-Only, Calculated Fields

In addition to a traditional, normalized RDBMs, a good single source of truth allows for building on top of this N-Dimensional "Knowledge Graph" with three additional types of Inline, Read-Only, Calculated Fields that are updated automatically at write-time, for any part of the database on which they depend.  In this way, the basic, normalized structure can be fleshed out with mult-dimensional, related meta data about each and every element in the graph, that automatically updates/maintains itself any time the underlying, normalized part of the data changes - allowing for an always internally consistent representation of literally any normalizable concept/system possible.

A DMZ of Knowledge

The knowledge graph then, which focuses exclusively on capturing facts in Tripples (i.e. 2 nodes + a directed-edge connecting them) in an extremely granular and precise format.  There can be N agents updating this graph - each completely unaware (and unaffected) by how many agents there are on the other side of the DMZ consuming the graph.

The benefit of the graph, is that like the translation between AI and Human, where we don't speak their language at all, then can still communicate their knowledge, unambiguously, by simply responding to any "question" or "query" that we put to it, with a Json blob of data containing facts that unambigously settle the question being asked.

It will not be a narrative description of the answer, because we already tried that.  We had all our agents (the people) write down everything in their head and publish that "language" to "the internet" and then we put AI "search engines" in front of all of those documents.  In this model, when we ask a question, the protocol is that the now increasingly intelligent "search agent" gets to return 10 links to us, where hopefully, somehwere in the first few links will be some text that may/may not answer our question.

We tried this from the early 90's through ~2015.  

When you ask google "When was the war of 1812" - it does not start parsing the pages of the internet to answer this question.  

This is still the essentially the model for the left side of Google, but increasingly, the actual answer to your question can be found directly, as a direct answer to the actual question asked - on the right side of the page.  The reason for this is that the right side of the page is an automated report, built from the Json blob of actual answers to the given question, which is the result from querying a Knowledge Graph.  

Putting it all together

With this powerful framework unlerlying all of the technology, the Single Source of Truth that lives "in the cloud" can be exported in pieces, so that each consumer needs only look at the part of the graph that relates to the part of the problem being addressed, and yet, with confidence that each small piece correctly relates to the whole, in the grand scheme of things.

This structural representation of the idea makes for a powerful framework on top of which to build cross-platform software that will tend to remain internally consistent, because most of the decisions are simply following their upstream knowledge graph.

Saturday, October 1, 2022

Any "No-Code App" IS (by definition) a Knowledge Graph

Knowledge graphs are pretty abstract mathematical objects, and training someone how to build a really good knowledge graph could potentially take months or years of training.

It turns out though, that when it comes to building a knowledge graph specifically to describe how a particular piece of software should work, if we simply build a prototype of the software we want, in any no  code tool on the planet, that can be exported to Json, then we will have designed a knowledge graph by definition.

In other words, If we were to use NoCodeXYZ.com to design a little prototype app that does exactly what we want, and then export it to a Json file - my-project.json, and then delete the original app on NoCodeXYZ.com, then the only place that the decisions about how that app worked still exist would be, by definition at this point, in that Json file.

This could be shown to be true (by definition) by Re-Importing the my-project.json file back into NoCodeXYZ.com.  If the resulting app does everything that we want, then we can infirm that 100% of the knowledge, about how the app should be have must have been embedded somewhere in that Json export.

Indeed, I would go a step further, and argue that given these underlying facts, we should always be able to build a Knowledge Graph from the Json export provided - and that if this knowledge graph is then shared with the "Actual" development team - most of the "Specification Documentation" will become entirely unnecessary, because, rather than constantly having to hard-code all of the actual rules into their code, in language, after language, after language - they can simply write their rules relative to the facts defined in the Json file available at design time.

This this in mind, we can have an English Specification Report (which is generated from the Knowledge Graph), and then code in 10 languages, which each also reference the knowledge graph. And then, when the decisions change (as they always do) we can simply update the knowledge graph and then regenerate the Specification Report and Json file - and all 10 languages are now back on the same page.

By contrast... if the code in those 10 languages was all just hard coded, following rules written down in an English specification document, when a decision changes, this is the traditional method of implementing said change:

1) A decision changes
2) A change request is made which describes (in English) what changed.  This change request is now out of synch with the original Design Specification, because the original spec is not usually updated/kept up to date as decisions change.
3) This change request gets scheduled into some subsequent sprint, when the various developers who implemented the original logic can be assigned to jump back into their code, and implement the change requested.
4) Once changed, they now submit the code to be tested and released.
5) IF any bugs are discovered, the process repeats and we go back to step 2.

Now, to be clear, this is not to say that the steps described above will never happen when the project has a knowledge graph, it's just that they will need to happen dramatically less often, because many/most of the decisions can simply reference the knowledge graph - which avoids tight coupling between those decisions, and the various technical contexts and environments in which that decision is expected to be implemented in.



Monday, September 5, 2022

Requirements Gathering ... and how to define a "thing"

Gathering requirements is one of the hardest and most important aspects of software development.


The world is full of ideas, and when one person tries to communicate an idea to another person, They’re stuck with having to pick a specific word to identify each thing that they’re talking about.


So when one person uses a word to describe a thing, if the person they’re speaking with is familiar with that word they will think of the same thing when they hear it. 


But if they’re not familiar with it, or have a different understanding, the result is that they are actually thinking about a different thing.


Same word, different thing.


This is a fundamental problem with language. Human language, spoken language, written language, programming languages, all languages.


A different approach


For the first time in history, though, we have a tool in the computer that allows us to take an entirely different approach.


In about the same amount of time that it would take to describe a thing, in 2022 we can simply create a digital instance of the thing being described. Just one.


The critical aspect of this one, digital example of the idea, is that the things being described exist (or don’t exist) within the digital model being built, regardless of what you call it.


Each person might actually call it something different, but within the model _it_ only exists once. And it’s relationships to the other things being discussed are also relationships that exist independently of what you call the various participants in that relationship.


In this way, answers to simple questions become possible, even trivial to answer – whereas, with a linguistic description of an idea there will always be ambiguity about whether two different words refer to the same thing or to two different things, making it difficult to even just count how many “things” are being discussed.  


Knowing how many things are involved is not typically an important piece of information. But… 


Being *able* to answer that question versus having absolutely no mechanism to even begin to answer such a question reflects a deep, deep flaw in literally any linguistic description of a complex idea.


Structure of the digital twin


The digital twin, or model, or single source of truth as we call it - is actually two separate things mixed into one.


In other words, in order to take a piece of language and extract from it a digital model of the ideas being discussed, there are two parts of that conversation.


In the first part of the conversation, we have to agree about what the structure of that model will be. Literally every single idea will have a slightly different structure.


In fact, the structure Of that model essentially act as a fingerprint for the idea.


Once we can agree on that model though, we can then begin to take the Content of the idea and enter it into that model.


But this is an iterative process, so anytime the contact doesn’t fit into the model, it prompts another discussion About and potentially changes to the structure of the model. The scheme of the idea.  


The last part of the process then, is to write a report against the single source of truth, which can reassemble the words back into a linguistic description of the idea.


What essential however, is that the linguistic description is not the source. It is a product that is produced from the source, which is a digital, multi dimensional representation of the idea itself. 


Thursday, September 1, 2022

Language is like a picture of a sculpture

 English might be an oil painting. German, a charcoal sketch. French, a watercolor. 


Each a slightly different representation of the same marble sculpture of a beautiful horse.


A Single Source of Truth, by contrast, is a scale replica of the statue itself.


Let’s examine the question of whether the statue is a horse, or whether it’s actually a mythical creature, the winged Pegasus.  


The first picture we look at, from the front might not show whether or not the statue has wings.


Another picture of the statue from the side might make it look like it has wings, and yet another picture from a slightly different perspective might make it clear that those wings actually belong to another animal behind the horse. 


Part of the problem with this model is that every time a new idea is introduced, more pictures get added to the pool of evidence. In other words, each opinion considered ends up being more and more subjective, as to which evidence is included, or not, when considering the question at hand.  


With a single source of truth, by contrast, the question of whether or not the horse has wings becomes a binary, easily verifiable fact. There is very, very little room for disagreement when we can all just look at the same horse.


In this case, the problem emerges because the statue is three-dimensional and the pictures are all two dimensional. Because of this, it requires many, many, many pictures, each only capturing part of the actual statue itself.  The problem is further amplified if we have to fall back to one-dimensional language to describe the statue.


If we can just create the statue itself then we don’t need many of them. We only need one. 


At the heart of the matter, ideas are multi-dimensional.  So trying to describe ideas with one dimensional language, two dimensional pictures, or even 3d sculptures is usually an exercise in futility.


But, with computers and literally just some simple tools, literally hailing from the 1960s and 70s, we can sketch up a digital twin of even complex, multi-dimensional ideas… effectively at the speed of speech.   


Not a description of the idea.  Not a picture of the sculpture.  But rather a scale replica of the thing itself - a digital twin of the abstract notions described; a model of the claims made, which unambiguously encodes the facts described.  


So, not a description of the idea, but a scale model, which digitally mirrors what was described originally. Like sheet music for the idea.


With this sculpture, this single source of truth in hand, the painter can paint the horse from any angle he chooses. The sketch artist can sketch him; the writer can describe him; and all of these different perspectives will tend to match each other because they’re all looking at the same scale model of the sculpture for motivation.


And if one of them adds wings to their horse, literally anyone should be able to just glance at the sculpture and say, definitely, that the painting does not accurately represent the horse. 


It should not be a debate. 


Without the sculpture in hand, by contrast, it’s always a debate, facilitated by adding more and more perspectives of the same thing. From some angles it will look like the horse has wings. From other angles, it won’t. And thus the debate rages on.


Let's change the conversation, by agreeing to shared context.  A shared model.  A shared set of facts.  


I.e. a Single Source of Truth