25-Comp-B10 Distributed Systems · May 2014
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
(a) Distributed vs. centralized systems — two advantages, two disadvantages. Advantage 1 — resource sharing and incremental scalability. A distributed system lets many independent machines' processing power, storage and peripherals be pooled and shared across an organization, and capacity can be grown incrementally by adding more machines rather than being permanently bounded by whatever ceiling a single centralized machine happens to have. Advantage 2 — fault tolerance and availability. Because functionality and data can be replicated across independent nodes, the failure of any one machine need not halt the whole system — other nodes continue serving, whereas a centralized system has a single point of failure: one crash stops everything. Disadvantage 1 — added complexity. A centralized system has one clock, one memory, and no partial failure or concurrent access from a network to reason about; a distributed system must additionally handle concurrency, partial failure (some nodes up, some down), the absence of a single global clock, and unpredictable message delay — all of which make correctness far harder to reason about and to test than a single-machine program. Disadvantage 2 — security exposure. Every message that crosses the network between nodes is a potential interception or tampering point, so a distributed system must actively defend an inherently open, shared network path between its components, whereas a centralized system's data mostly stays inside one machine's protected memory and never has to cross an untrusted wire at all.
(b) Client-server architecture of a major Internet application. In the client-server model, a server is a program that runs continuously (or on demand), listens on a well-known network address/port, and provides a defined service by responding to requests; a client is a program that initiates communication by sending a request to a server and consuming the reply, typically on behalf of an interactive user. The relationship is asymmetric: the server is passive (waits to be asked) and typically shared by many concurrent clients, while each client instance is usually private to one user or one task and initiates the interaction. The two communicate over a network using a request-reply protocol built on top of transport-layer sockets (TCP or UDP): the client marshals its request into a message, sends it to the server's address, and blocks (or polls asynchronously) until the reply message arrives.
The same skeleton generalizes to email and ftp: an email client speaks SMTP to a mail-submission server and IMAP/POP3 to a mailbox server to fetch messages; an ftp client opens a control connection to negotiate commands and a separate data connection for the actual file transfer. In every case, the server exposes a stable, addressable service and the client is the one that must know where to find it and initiate contact.
(c) Resources shared efficiently in distributed systems. Software resources: (1) a shared database or file-server volume — e.g. a company's order-management database is hosted once and accessed concurrently by many client applications over the network, with the DBMS serializing/coordinating concurrent updates so clients never see a fragmented view; (2) a shared network/print service — a print server maintains one spool queue and one physical printer driver that many client workstations submit jobs to, rather than every desktop needing its own driver and direct cable to a printer. Hardware resources: (1) high-capacity storage (a NAS/SAN array) — many clients mount the same network volume so expensive, high-reliability (RAID-protected) disks are amortized across an organization instead of duplicated per desktop; (2) specialized compute hardware, e.g. a GPU/compute cluster — scientific or rendering jobs are submitted from many client machines to a shared cluster of accelerators, which is far more cost-effective than equipping every workstation with the same capability it uses only occasionally. In each case the resource is efficient to share precisely because individual clients need it only intermittently, while the server keeps it centrally available and multiplexes access across the community of users.
(d) Administrative scalability. Administrative scalability is the ability of a system to keep working well as it grows to span many separately administered organizations or domains, not merely to handle more nodes or users within a single, unified administrative authority. It is often the hardest kind of scalability to achieve because it is fundamentally a governance problem rather than a raw capacity problem: each administrative domain has its own security policy, its own naming authority, its own resource-allocation rules, and possibly conflicting requirements, and no single organization can simply mandate a change across all of them the way an in-house IT department can push an update across its own machines. A classic example is the Domain Name System, which achieves administrative scalability by delegating naming authority hierarchically — each domain administers its own subtree of names autonomously — rather than one central body maintaining a single flat name table for the whole Internet; this delegation is precisely what lets DNS span millions of mutually independent organizations without requiring them to trust or coordinate with each other directly. The difficulty shows up whenever a genuinely global change is needed (e.g. rolling out a new protocol version, or reconciling incompatible privacy or trust policies between two organizations) since there is no single administrative arbiter empowered to force adoption everywhere at once — adoption instead has to happen gradually, domain by autonomous domain, and any solution must tolerate the domains that have not yet, or never will, come along.