NivaarExam PrepOfficial exam papers ↗

25-Comp-B10 Distributed Systems · Undated paper

Question 1 of 6: Characteristics of Distributed Systems

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

17-Comp-B10 Distributed Systems — National Exams, May 2019. 3 hours, closed book, Casio/Sharp approved calculator only. Candidates were instructed to answer any five of the six questions, only the first five as they appear in the answer book marked, with most questions requiring an essay-format answer; all six are answered below as a complete study resource.

Reference texts: Coulouris, Dollimore, Kindberg & Blair, Distributed Systems: Concepts and Design (5th ed.) — system models, characterization of distributed systems and openness (ch. 1–2), networking and internetworking, TCP/IP (ch. 3), interprocess communication and remote invocation, RPC (ch. 4–5), operating system support (ch. 6–7), security (ch. 11), distributed file systems (ch. 12).

Question 1: Characteristics of Distributed Systems (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Three common challenges of building distributed systems. Heterogeneity. A distributed system must span different networks, hardware architectures, operating systems, programming languages and vendor implementations; it copes by standardizing on network protocols (IP/TCP), middleware (RPC/RMI, Web services) and neutral data-representation formats so components built on different platforms can still interoperate. Security. Resources shared across an open network are exposed to eavesdropping, tampering, impersonation and denial of service by parties the system never explicitly trusted; encryption, authentication and secure channels are needed to close a gap that simply does not exist in a single, physically-secured machine. Failure handling. In a single machine, a crash normally stops the whole program; in a distributed system, individual components (a client, a server, the network) can fail independently while the rest of the system keeps running, so partial failure — not total failure — is the norm, and every remote interaction must be designed with timeouts, retries and redundancy rather than the certainty a local call enjoys.

(b) Transparency. Transparency is the property that hides the distribution of a system's components — their physical location, the mechanism used to access them, the fact that they may fail or be replicated — so that the system appears to users and application programmers as a single, coherent whole rather than a collection of independent, networked parts. It matters because it is what lets programmers reason about and build distributed applications using largely the same concepts as a single-machine program: without it, every piece of client code would need to hard-code knowledge of exactly where each resource lives, how to reach it, and how to work around every possible partial failure, which does not scale as the system grows or changes. A concrete example is location transparency: a client opens a file by its logical name (e.g. through NFS's uniform file-naming interface) with no idea, and no need to know, which physical server actually stores it — the file can even be moved to a different server later without the client's code changing at all.

(c) Open distributed systems and the benefits of openness. An open distributed system is one whose key interfaces are published according to a well-defined, publicly available specification (rather than kept as a proprietary vendor secret), so that independently-written software from different vendors can interoperate with it, extend it, or be substituted into it. Openness delivers: interoperability — components from different vendors, built to the same published interface, can work together (e.g. any standards-compliant Web browser can talk to any standards-compliant Web server); portability — an application written against the open interface can run on different underlying implementations/platforms with little or no change; and extensibility — because the specification is public, new services can be added or existing ones re-implemented (even by third parties) without breaking existing clients, avoiding vendor lock-in to any single supplier's roadmap.

(d) The Web as resource sharing / client-server; HTML, URLs and HTTP as core browsing technologies. The World Wide Web is a canonical resource-sharing system: a server (a Web server process, potentially one of many replicas behind a load balancer) hosts resources — HTML pages, images, data — each named by a URL, and a client (a browser) sends a request naming the resource it wants and receives it in a reply, with no need to know which disk or file path actually stores that resource. Many independent clients share the very same resource concurrently, and the server (or an intermediate cache/proxy) handles the concurrency and scaling this implies; the same program can even play both roles — a Web server acting as a client of a database server, or a search engine's crawler acting as a Web client.

HTML — advantages: a simple, platform-independent markup language that any browser can render, whose embedded hyperlinks let any document reference any other resource anywhere on the Web, which is exactly what makes the Web an open, extensible information space that anyone can add to. Disadvantages: it describes how content is presented rather than what it means, so the data it carries is hard for programs to process (hence XML/JSON for machine exchange), and a static page has no interaction of its own — richer behaviour needs scripts, forms and other technologies layered on top. URLs — advantages: one uniform naming scheme (scheme + DNS host name + path, e.g. http://www.example.ca/index.html) that identifies a resource of any type and is by itself enough to locate and fetch it from anywhere. Disadvantages: a URL embeds the server's host name and the path on that server, so it is location-dependent — if a resource moves or is deleted, every link to it breaks (dangling links, “404 Not Found”), and there is no referential integrity between documents. HTTP — advantages: a simple, stateless request-reply protocol with a small set of methods (GET, POST, PUT, ...) and content-type negotiation, so any kind of data can be transferred; statelessness means any request can be served by any cache, proxy or replica, which is what lets the Web scale. Disadvantages: statelessness must be worked around (cookies/tokens) for session-oriented interaction; early versions opened a new TCP connection for every resource fetched, adding latency (mitigated by persistent connections and HTTP/2 multiplexing); and plain HTTP gives no confidentiality or integrity of its own, needing TLS (HTTPS) on top.

← Paper overview